用单张图编辑2D图像,实现动态3D场景的精准可控局部修改。
CTRL-D: Controllable Dynamic 3D Scene Editing with Personalized 2D Diffusion
- 通过微调InstructPix2Pix模型,从一张参考图学习编辑能力。
- 两阶段优化使编辑结果在时间上保持一致,加速收敛并提升质量。
- 无需跟踪目标区域,适合需要精细控制的3D场景编辑任务。
神经辐射场和3D高斯泼溅等3D表示技术显著提升了真实感场景建模与新视角合成效果。然而,动态3D场景中实现可控且一致的编辑仍具挑战,现有方法受限于编辑框架,导致编辑不一致、控制力弱。本文提出一种新框架:先微调InstructPix2Pix模型,再基于可变形3D高斯进行两阶段优化。微调使模型能从单张编辑参考图中“学习”编辑能力,将复杂的动态场景编辑转化为简单的2D图像编辑。通过直接学习编辑区域与风格,无需追踪目标区域即可实现一致且精确的局部编辑。随后,利用设计的已编辑图像缓冲区加速收敛并提升时序一致性。相比当前最优方法,本方案提供更灵活、可控的局部编辑能力,生成高质量且一致的结果。
原文摘要 · Abstract (English)
Recent advances in 3D representations, such as Neural Radiance Fields and 3D Gaussian Splatting, have greatly improved realistic scene modeling and novel-view synthesis. However, achieving controllable and consistent editing in dynamic 3D scenes remains a significant challenge. Previous work is largely constrained by its editing backbones, resulting in inconsistent edits and limited controllability. In our work, we introduce a novel framework that first fine-tunes the InstructPix2Pix model, followed by a two-stage optimization of the scene based on deformable 3D Gaussians. Our fine-tuning enables the model to "learn" the editing ability from a single edited reference image, transforming the complex task of dynamic scene editing into a simple 2D image editing process. By directly learning editing regions and styles from the reference, our approach enables consistent and precise local edits without the need for tracking desired editing regions, effectively addressing key challenges in dynamic scene editing. Then, our two-stage optimization progressively edits the trained dynamic scene, using a designed edited image buffer to accelerate convergence and improve temporal consistency. Compared to state-of-the-art methods, our approach offers more flexible and controllable local scene editing, achieving high-quality and consistent results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。