arXiv:2508.13797cs.GRcs.CV2025-08International Conf…被引 5

用草图精准编辑带大视角变化的3D视频场景,保持结构一致。

Sketch3DVE: Sketch-based 3D-Aware Scene Video Editing

  • 以草图控制3D几何,结合点云与深度图实现精确编辑
  • 通过3D-aware掩码传播和扩散模型生成逼真新视图内容
  • 支持大幅相机旋转/缩放下的稳定编辑,适合影视后期应用

现有视频编辑方法在风格迁移或外观修改上表现良好,但对带显著视角变化(如大幅相机旋转或缩放)的3D场景结构内容编辑仍具挑战。主要难点包括生成与原视频一致的新视角内容、保留未编辑区域、以及将稀疏2D输入转化为真实3D视频输出。为此,我们提出Sketch3DVE,一种基于草图的3D感知视频编辑方法,支持在大视角变化下进行细节化局部编辑。针对稀疏输入问题,先用图像编辑生成首帧结果并传播至后续帧;采用草图作为交互工具实现精确几何控制,亦兼容掩码编辑。为应对视角变化,我们通过密集立体匹配估计输入视频的点云与相机参数,并提出基于深度图表示新编辑部分的点云编辑方法,有效对齐原始3D场景。为无缝融合新内容并保留未编辑区域特征,引入3D-aware掩码传播策略,并利用视频扩散模型生成逼真编辑结果。大量实验表明,Sketch3DVE在视频编辑任务中具有明显优势。

原文摘要 · Abstract (English)

Recent video editing methods achieve attractive results in style transfer or appearance modification. However, editing the structural content of 3D scenes in videos remains challenging, particularly when dealing with significant viewpoint changes, such as large camera rotations or zooms. Key challenges include generating novel view content that remains consistent with the original video, preserving unedited regions, and translating sparse 2D inputs into realistic 3D video outputs. To address these issues, we propose Sketch3DVE, a sketch-based 3D-aware video editing method to enable detailed local manipulation of videos with significant viewpoint changes. To solve the challenge posed by sparse inputs, we employ image editing methods to generate edited results for the first frame, which are then propagated to the remaining frames of the video. We utilize sketching as an interaction tool for precise geometry control, while other mask-based image editing methods are also supported. To handle viewpoint changes, we perform a detailed analysis and manipulation of the 3D information in the video. Specifically, we utilize a dense stereo method to estimate a point cloud and the camera parameters of the input video. We then propose a point cloud editing approach that uses depth maps to represent the 3D geometry of newly edited components, aligning them effectively with the original 3D scene. To seamlessly merge the newly edited content with the original video while preserving the features of unedited regions, we introduce a 3D-aware mask propagation strategy and employ a video diffusion model to produce realistic edited videos. Extensive experiments demonstrate the superiority of Sketch3DVE in video editing. Homepage and code: http://http://geometrylearning.com/Sketch3DVE/

3D视频编辑草图交互扩散模型点云建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。