arXiv:2410.12836cs.GRcs.AI2024-10ICLR被引 15

用自然语言一键编辑3D房间布局,支持增删改等六种操作。

EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing

  • 通过大模型解析指令,用扩散模型生成新场景。
  • 在83000对数据上训练,六类编辑任务均优于基线。
  • 适合游戏、VR/AR开发人员快速搭建3D场景。

针对专业3D软件学习成本高、资产管理耗时的问题,语言引导的3D场景编辑具有重要应用价值。现有方法或需人工干预,或仅能修改外观而无法调整布局。为此,我们提出EditRoom,一个统一框架,可通过自然语言指令无须人工干预完成多种布局编辑。该框架利用大语言模型(LLM)进行指令规划,并采用基于扩散的方法生成目标场景,支持旋转、平移、缩放、替换、添加和删除共六类编辑操作。为解决语言引导3D场景编辑数据匮乏问题,我们设计了自动化数据增强流程,构建了包含83,000个编辑对的EditRoom-DB大规模数据集,用于训练与评估。实验表明,我们的方法在所有指标上均显著优于基线,展现出更高的准确性和场景一致性。

原文摘要 · Abstract (English)

Given the steep learning curve of professional 3D software and the time-consuming process of managing large 3D assets, language-guided 3D scene editing has significant potential in fields such as virtual reality, augmented reality, and gaming. However, recent approaches to language-guided 3D scene editing either require manual interventions or focus only on appearance modifications without supporting comprehensive scene layout changes. In response, we propose EditRoom, a unified framework capable of executing a variety of layout edits through natural language commands, without requiring manual intervention. Specifically, EditRoom leverages Large Language Models (LLMs) for command planning and generates target scenes using a diffusion-based method, enabling six types of edits: rotate, translate, scale, replace, add, and remove. To address the lack of data for language-guided 3D scene editing, we have developed an automatic pipeline to augment existing 3D scene synthesis datasets and introduced EditRoom-DB, a large-scale dataset with 83k editing pairs, for training and evaluation. Our experiments demonstrate that our approach consistently outperforms other baselines across all metrics, indicating higher accuracy and coherence in language-guided scene layout editing.

3D生成语言控制扩散模型场景编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。