arXiv:2602.09153cs.ROcs.AI2026-02被引 24

用自然语言生成逼真室内场景,提升机器人训练真实性

SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes

  • 分阶段由视觉语言代理协作构建场景
  • 生成物体数达基线3-6倍,碰撞率低于2%
  • 适合机器人仿真与策略评估研究者使用

仿真已成为大规模训练和评估家用机器人的重要工具,但现有环境难以还原真实室内空间的多样性和物理复杂性。当前场景生成方法产生的房间家具稀疏,缺乏密集杂物、可动家具及物理属性,难以支持机器人操作任务。我们提出SceneSmith,一种层级化智能体框架,能根据自然语言提示生成仿真就绪的室内环境。该框架通过设计、批评与协调三类视觉语言代理,在建筑布局、家具摆放、小物件分布等阶段逐级协作生成场景。系统融合文本到3D生成、数据集检索与物理属性估计,实现静态物体生成、可动部件获取与物理特性预测。SceneSmith生成的物体数量为先前方法的3至6倍,物体间碰撞率低于2%,96%的物体在物理模拟中保持稳定。205名用户参与的评估显示,其生成场景在真实感上达到92%平均得分,提示忠实度达91%胜率。进一步实验证明,这些环境可直接用于端到端机器人策略自动评估流程。

原文摘要 · Abstract (English)

Simulation has become a key tool for training and evaluating home robots at scale, yet existing environments fail to capture the diversity and physical complexity of real indoor spaces. Current scene synthesis methods produce sparsely furnished rooms that lack the dense clutter, articulated furniture, and physical properties essential for robotic manipulation. We introduce SceneSmith, a hierarchical agentic framework that generates simulation-ready indoor environments from natural language prompts. SceneSmith constructs scenes through successive stages$\unicode{x2013}$from architectural layout to furniture placement to small object population$\unicode{x2013}$each implemented as an interaction among VLM agents: designer, critic, and orchestrator. The framework tightly integrates asset generation through text-to-3D synthesis for static objects, dataset retrieval for articulated objects, and physical property estimation. SceneSmith generates 3-6x more objects than prior methods, with <2% inter-object collisions and 96% of objects remaining stable under physics simulation. In a user study with 205 participants, it achieves 92% average realism and 91% average prompt faithfulness win rates against baselines. We further demonstrate that these environments can be used in an end-to-end pipeline for automatic robot policy evaluation.

场景生成机器人仿真多智能体文本到3D

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。