用双场表示提升文本驱动3D场景编辑的清晰度与稳定性
DualNeRF: Text-Driven 3D Scene Editing via Dual-Field Representation
- 引入双场表示,保留原始场景特征作为背景维护指导
- 模拟退火策略解决优化过程中的局部最优问题
- 结合CLIP一致性指标过滤低质量编辑,适合高质量3D生成任务
近期,去噪扩散模型在2D图像生成与编辑中取得显著进展。Instruct-NeRF2NeRF(IN2N)通过“迭代数据集更新”(IDU)策略将扩散模型的成功引入3D场景编辑。尽管结果引人注目,但IN2N存在背景模糊和陷入局部最优的问题:前者源于缺乏对背景的有效引导,后者则由图像编辑与NeRF训练在IDU过程中相互干扰所致。本文提出DualNeRF以解决上述问题。我们设计了一种双场表示,用于保留原始场景特征,并在IDU过程中将其作为额外指导以维持背景一致性。此外,将模拟退火策略嵌入IDU,赋予模型跳出局部最优的能力。同时,采用基于CLIP的一致性指示器,通过过滤低质量编辑进一步提升生成效果。大量实验表明,本方法在定性和定量评估上均优于现有方法。
原文摘要 · Abstract (English)
Recently, denoising diffusion models have achieved promising results in 2D image generation and editing. Instruct-NeRF2NeRF (IN2N) introduces the success of diffusion into 3D scene editing through an "Iterative dataset update" (IDU) strategy. Though achieving fascinating results, IN2N suffers from problems of blurry backgrounds and trapping in local optima. The first problem is caused by IN2N's lack of efficient guidance for background maintenance, while the second stems from the interaction between image editing and NeRF training during IDU. In this work, we introduce DualNeRF to deal with these problems. We propose a dual-field representation to preserve features of the original scene and utilize them as additional guidance to the model for background maintenance during IDU. Moreover, a simulated annealing strategy is embedded into IDU to endow our model with the power of addressing local optima issues. A CLIP-based consistency indicator is used to further improve the editing quality by filtering out low-quality edits. Extensive experiments demonstrate that our method outperforms previous methods both qualitatively and quantitatively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。