用大模型推理生成更符合指令的3D场景,解决物体摆放不准的问题。
Text-to-Scene with Large Reasoning Models
- 利用大推理模型理解文本指令并检索带属性的物体
- 结合隐式与显式布局约束,碰撞感知调整物体位置
- 适合需要高精度场景还原的研究者和开发者
提示驱动的场景合成允许用户从文本描述生成完整的3D环境。现有文本到场景方法在复杂几何结构和物体变换上表现不佳,且对复杂指令遵循能力弱。我们提出Reason-3D,一个由大推理模型(LRMs)驱动的文本到场景模型。Reason-3D通过包含物理、功能和上下文属性的图像描述进行物体检索,随后基于隐式与显式布局约束放置选定物体,并通过碰撞感知的空间推理优化其位置。在从简单到复杂的室内配置指令上评估,Reason-3D在人类评分的视觉保真度、约束遵循度和资产检索质量上显著优于先前方法。本工作不仅推动了文本到场景生成领域的发展,还展示了现代大推理模型的先进空间推理能力。此外,我们开源代码库,以促进基于大推理模型的物体检索与定位研究。
原文摘要 · Abstract (English)
Prompt-driven scene synthesis allows users to generate complete 3D environments from textual descriptions. Current text-to-scene methods often struggle with complex geometries and object transformations, and tend to show weak adherence to complex instructions. We address these limitations by introducing Reason-3D, a text-to-scene model powered by large reasoning models (LRMs). Reason-3D integrates object retrieval using captions covering physical, functional, and contextual attributes. Reason-3D then places the selected objects based on implicit and explicit layout constraints, and refines their positions with collision-aware spatial reasoning. Evaluated on instructions ranging from simple to complex indoor configurations, Reason-3D significantly outperforms previous methods in human-rated visual fidelity, adherence to constraints, and asset retrieval quality. Beyond its contribution to the field of text-to-scene generation, our work showcases the advanced spatial reasoning abilities of modern LRMs. Additionally, we release the codebase to further the research in object retrieval and placement with LRMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。