arXiv:2602.14968cs.ROcs.AI2026-02被引 6

用大模型+物理引擎生成高复杂度真实物理场景,提升机器人仿真数据质量。

PhyScensis: Physics-Augmented LLM Agents for Complex Physical Scene Arrangement

  • 大模型迭代提出物体位置与物理关系,物理引擎验证并反馈修正。
  • 生成场景物体密度达数十个,稳定性与空间关系控制精度显著提升。
  • 适合需要复杂物理交互的机器人操控仿真,尤其桌面/货架场景。

自动生成可交互的3D环境对大规模机器人仿真数据收集至关重要。现有工作多关注3D资产摆放,却常忽略物体间的物理关系(如接触、支撑、平衡、容纳),这些是构建复杂真实操作场景(如桌面布置、货架整理、箱体堆叠)的关键。相比传统3D布局生成,复杂物理场景面临更高物体密度与复杂度(如小货架可容纳数十本书)、更丰富的支撑关系与紧凑布局,以及精确建模空间位置与物理属性的需求。为此,我们提出PhyScensis——一种基于大模型代理与物理引擎的框架,可生成高复杂度且物理解释合理的场景配置。该框架包含三部分:大模型代理迭代提出带空间与物理谓词的资产;求解器结合物理引擎将谓词实现为3D场景;求解器反馈驱动代理优化与丰富配置。此外,通过概率编程控制稳定性,并结合启发式方法协同调节稳定性和空间关系,实现对细粒度文本描述与数值参数(如相对位置、场景稳定性)的强可控性。实验表明,该方法在场景复杂度、视觉质量与物理准确性上均优于现有方法,为机器人操作提供了统一的复杂物理场景生成管道。

原文摘要 · Abstract (English)

Automatically generating interactive 3D environments is crucial for scaling up robotic data collection in simulation. While prior work has primarily focused on 3D asset placement, it often overlooks the physical relationships between objects (e.g., contact, support, balance, and containment), which are essential for creating complex and realistic manipulation scenarios such as tabletop arrangements, shelf organization, or box packing. Compared to classical 3D layout generation, producing complex physical scenes introduces additional challenges: (a) higher object density and complexity (e.g., a small shelf may hold dozens of books), (b) richer supporting relationships and compact spatial layouts, and (c) the need to accurately model both spatial placement and physical properties. To address these challenges, we propose PhyScensis, an LLM agent-based framework powered by a physics engine, to produce physically plausible scene configurations with high complexity. Specifically, our framework consists of three main components: an LLM agent iteratively proposes assets with spatial and physical predicates; a solver, equipped with a physics engine, realizes these predicates into a 3D scene; and feedback from the solver informs the agent to refine and enrich the configuration. Moreover, our framework preserves strong controllability over fine-grained textual descriptions and numerical parameters (e.g., relative positions, scene stability), enabled through probabilistic programming for stability and a complementary heuristic that jointly regulates stability and spatial relations. Experimental results show that our method outperforms prior approaches in scene complexity, visual quality, and physical accuracy, offering a unified pipeline for generating complex physical scene layouts for robotic manipulation.

物理仿真大模型代理场景生成机器人训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。