让2D编辑器学会保持3D场景多视角一致,提升3D编辑质量。
DisCo3D: Distilling Multi-View Consistency for 3D Scene Editing
- 用多视角输入训练3D生成器,再将一致性知识蒸馏到2D编辑器中。
- 在多个数据集上实现稳定多视角一致,编辑质量优于现有方法。
- 适合需要高质量3D场景编辑的从业者,如游戏与影视制作。
尽管扩散模型在2D图像生成与编辑方面取得显著进展,将其能力扩展至3D编辑仍面临挑战,尤其在维持多视角一致性方面。传统方法通常基于单一编辑视角进行迭代优化,但易导致收敛缓慢和跨视角不一致带来的模糊伪影。近期方法通过传播2D编辑注意力特征提升效率,但在复杂场景中仍存在细粒度不一致与失败模式,源于约束不足。为此,我们提出DisCo3D,一种将3D一致性先验蒸馏至2D编辑器的新框架。首先使用多视角输入微调3D生成器以适应场景,随后通过一致性蒸馏训练2D编辑器。最终,编辑后的多视角输出通过高斯溅射优化为3D表示。实验表明,DisCo3D实现了稳定的多视角一致性,在编辑质量上超越现有最优方法。
原文摘要 · Abstract (English)
While diffusion models have demonstrated remarkable progress in 2D image generation and editing, extending these capabilities to 3D editing remains challenging, particularly in maintaining multi-view consistency. Classical approaches typically update 3D representations through iterative refinement based on a single editing view. However, these methods often suffer from slow convergence and blurry artifacts caused by cross-view inconsistencies. Recent methods improve efficiency by propagating 2D editing attention features, yet still exhibit fine-grained inconsistencies and failure modes in complex scenes due to insufficient constraints. To address this, we propose \textbf{DisCo3D}, a novel framework that distills 3D consistency priors into a 2D editor. Our method first fine-tunes a 3D generator using multi-view inputs for scene adaptation, then trains a 2D editor through consistency distillation. The edited multi-view outputs are finally optimized into 3D representations via Gaussian Splatting. Experimental results show DisCo3D achieves stable multi-view consistency and outperforms state-of-the-art methods in editing quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。