用2D扩散模型精准定位3D局部编辑区域,实现高效一致的细节优化。
Mastering Regional 3DGS: Locating, Initializing, and Editing with Diverse 2D Priors
- 结合2D扩散模型与逆渲染,从多视角识别并定位3D编辑区域。
- 通过2D基础模型预测深度图,初始化粗略3DGS,实现视图一致性。
- 支持迭代编辑,提升结构与纹理一致性,速度最快达4倍加速。
许多3D场景编辑任务聚焦于局部区域而非整体场景,除风格迁移等全局应用外,3D高斯溅射(3DGS)通过一系列高斯表示,使局部编辑成为可能,实现对场景特定区域的精细控制;然而,3D语义解析性能常落后于2D,导致3D空间中的精准操作困难,影响编辑保真度。为此,本文利用2D扩散编辑在各视角中准确识别修改区域,通过逆渲染实现3D定位,再基于2D基础模型预测的深度图,优化前视图并初始化具一致视角和近似形状的粗略3DGS,支持迭代、视图一致的编辑流程,逐步增强结构细节与纹理,确保多视角连贯性。实验表明,该方法达到当前最优性能,同时实现最高4倍的速度提升,为3D场景局部编辑提供更高效、有效的解决方案。
原文摘要 · Abstract (English)
Many 3D scene editing tasks focus on modifying local regions rather than the entire scene, except for some global applications like style transfer, and in the context of 3D Gaussian Splatting (3DGS), where scenes are represented by a series of Gaussians, this structure allows for precise regional edits, offering enhanced control over specific areas of the scene; however, the challenge lies in the fact that 3D semantic parsing often underperforms compared to its 2D counterpart, making targeted manipulations within 3D spaces more difficult and limiting the fidelity of edits, which we address by leveraging 2D diffusion editing to accurately identify modification regions in each view, followed by inverse rendering for 3D localization, then refining the frontal view and initializing a coarse 3DGS with consistent views and approximate shapes derived from depth maps predicted by a 2D foundation model, thereby supporting an iterative, view-consistent editing process that gradually enhances structural details and textures to ensure coherence across perspectives. Experiments demonstrate that our method achieves state-of-the-art performance while delivering up to a $4\times$ speedup, providing a more efficient and effective approach to 3D scene local editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。