arXiv:2511.23172cs.CV2025-11AAAI被引 13

用视频生成模型的时序一致性,一键实现多视角一致的3D编辑。

Fast Multi-view Consistent 3D Editing with Video Priors

  • 用预训练视频模型生成其他视角的编辑结果,一次前向传播完成多视图编辑。
  • 单次前向传播即达成高质量3D编辑,速度比现有方法快数倍。
  • 适合需要快速生成一致多视角3D内容的研究者与创作者。

文本驱动的3D编辑允许用户通过文本指令轻松修改3D物体或场景。由于缺乏多视图一致性先验,现有方法通常依赖2D生成或编辑模型逐视图处理,再通过迭代2D-3D-2D更新,不仅耗时,且因不同视图编辑信号在迭代中平均化,易导致结果过平滑。本文提出基于生成式视频先验的3D编辑方法(ViP3DE),利用预训练视频生成模型的时序一致性先验,在单次前向传播中实现多视图一致的3D编辑。核心思路是将视频生成模型条件于单个已编辑视图,直接生成其他一致的编辑视图用于3D更新。由于3D更新需编辑视图与特定相机位姿配对,我们提出保持运动的噪声融合方法,使视频模型能在预设相机位姿下生成编辑视图。此外,引入几何感知去噪机制,将3D几何先验融入视频模型,进一步增强多视图一致性。大量实验表明,所提方法在单次前向传播下即可实现高质量3D编辑结果,显著优于现有方法在编辑质量和速度上的表现。

原文摘要 · Abstract (English)

Text-driven 3D editing enables user-friendly 3D object or scene editing with text instructions. Due to the lack of multi-view consistency priors, existing methods typically resort to employing 2D generation or editing models to process each view individually, followed by iterative 2D-3D-2D updating. However, these methods are not only time-consuming but also prone to over-smoothed results because the different editing signals gathered from different views are averaged during the iterative process. In this paper, we propose generative Video Prior based 3D Editing (ViP3DE) to employ the temporal consistency priors from pre-trained video generation models for multi-view consistent 3D editing in a single forward pass. Our key insight is to condition the video generation model on a single edited view to generate other consistent edited views for 3D updating directly, thereby bypassing the iterative editing paradigm. Since 3D updating requires edited views to be paired with specific camera poses, we propose motion-preserved noise blending for the video model to generate edited views at predefined camera poses. In addition, we introduce geometry-aware denoising to further enhance multi-view consistency by integrating 3D geometric priors into video models. Extensive experiments demonstrate that our proposed ViP3DE can achieve high-quality 3D editing results even within a single forward pass, significantly outperforming existing methods in both editing quality and speed.

3D编辑视频先验多视角一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。