用户拖拽即可精准编辑3D场景,支持多视角一致的局部修改。
DragScene: Interactive 3D Scene Editing with Single-view Drag Instructions
- 通过参考视图的隐空间优化生成2D编辑,实现直观拖拽操作。
- 利用点云重建粗略3D结构,确保多视角编辑一致性。
- 兼容多种3D表示,适合需要交互式编辑的创作者和设计师。
3D编辑在基于各类指令修改场景方面展现出强大能力。然而,现有方法难以实现直观、局部化的编辑,例如仅让特定花朵绽放。拖拽式编辑在图像编辑中表现出色,可通过直接操作替代模糊的文字指令。但将此类方法扩展至3D场景面临多视角不一致的重大挑战。为此,我们提出DragScene,一个将拖拽式编辑与多种3D表示相结合的框架。首先,在参考视图上进行隐空间优化,根据用户指令生成2D编辑结果。随后,利用点基表示从参考视图重建粗略的3D几何线索,捕捉编辑区域的细节。接着,将编辑后视图的隐空间表示映射至这些3D线索,指导其他视图的隐空间优化。该过程确保编辑在多视角间无缝传播并保持一致性。最后,从编辑后的多视图图像重建目标3D场景。大量实验表明,DragScene实现了对3D场景的精确、灵活的拖拽式编辑,且在多种3D表示中具有广泛适用性。
原文摘要 · Abstract (English)
3D editing has shown remarkable capability in editing scenes based on various instructions. However, existing methods struggle with achieving intuitive, localized editing, such as selectively making flowers blossom. Drag-style editing has shown exceptional capability to edit images with direct manipulation instead of ambiguous text commands. Nevertheless, extending drag-based editing to 3D scenes presents substantial challenges due to multi-view inconsistency. To this end, we introduce DragScene, a framework that integrates drag-style editing with diverse 3D representations. First, latent optimization is performed on a reference view to generate 2D edits based on user instructions. Subsequently, coarse 3D clues are reconstructed from the reference view using a point-based representation to capture the geometric details of the edits. The latent representation of the edited view is then mapped to these 3D clues, guiding the latent optimization of other views. This process ensures that edits are propagated seamlessly across multiple views, maintaining multi-view consistency. Finally, the target 3D scene is reconstructed from the edited multi-view images. Extensive experiments demonstrate that DragScene facilitates precise and flexible drag-style editing of 3D scenes, supporting broad applicability across diverse 3D representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。