arXiv:2504.15049cs.CV2025-04ICCV被引 6

用语言指令智能编辑真实3D扫描场景,兼顾物理合理与常识。

ScanEdit: Hierarchically-Guided Functional 3D Scan Editing

  • 构建分层场景图,实现复杂物体间的高效关联编辑。
  • 结合大模型推理与物理约束,生成符合常理的场景布局。
  • 支持多样语言指令,适合交互式3D内容创作应用。

随着3D捕获技术的快速发展和3D数据的激增,高效的3D场景编辑对图形学应用至关重要。本文提出ScanEdit,一种基于指令驱动的功能性3D扫描编辑方法。针对复杂、相互依赖的物体集合,我们提出分层引导的建模方式:首先将3D扫描分解为对象实例,并构建分层场景图以实现高效可扩展的编辑;随后利用大语言模型(LLMs)的能力,将高层语言指令转化为作用于场景图的分层操作命令;最后,ScanEdit融合基于LLM的指导与显式物理约束,生成在物体布局上既符合物理规律又符合常识的真实场景。在广泛的实验评估中,ScanEdit超越现有最优方法,在多种真实场景和输入指令下均表现出色。

原文摘要 · Abstract (English)

With the fast pace of 3D capture technology and resulting abundance of 3D data, effective 3D scene editing becomes essential for a variety of graphics applications. In this work we present ScanEdit, an instruction-driven method for functional editing of complex, real-world 3D scans. To model large and interdependent sets of ob- jectswe propose a hierarchically-guided approach. Given a 3D scan decomposed into its object instances, we first construct a hierarchical scene graph representation to enable effective, tractable editing. We then leverage reason- ing capabilities of Large Language Models (LLMs) and translate high-level language instructions into actionable commands applied hierarchically to the scene graph. Fi- nally, ScanEdit integrates LLM-based guidance with ex- plicit physical constraints and generates realistic scenes where object arrangements obey both physics and common sense. In our extensive experimental evaluation ScanEdit outperforms state of the art and demonstrates excellent re- sults for a variety of real-world scenes and input instruc- tions.

3D编辑语言指令场景生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。