arXiv:2606.01649cs.CV2026-06中稿 · ICML

让机器人桌场景生成更真实,自动避撞且符合物理规律。

PhyScene3D: Physically Consistent Interactive 3D Tabletop Scene Generation

论文配图:PhyScene3D: Physically Consistent Interactive 3D Tabletop Scene Generation
图 1 · 摘自论文原文
  • 模拟人类搭建过程,分步放置物体并约束位置。
  • 生成场景碰撞率比人工数据低40%,物理合理性显著提升。
  • 适合机器人交互、仿真训练等需要真实物理环境的场景

生成符合物理规律的3D桌面场景是交互式通用机器人学习中的基础但研究不足的问题,挑战源于密集物体层级与不规则可用性。本文定义的交互场景指可直接加载至物理模拟器的无碰撞、物理有效环境。现有方法从解耦符号求解到端到端回归模型,常因误差传播或对含大量物理违规的噪声标注过拟合。为此,提出PhyScene3D框架,将生成重构为类人建构过程。提出的认知拓扑推理链(CTRC)将场景合成分解为序列化、锚点条件化的步骤,采用基于3D AABB的放置策略,引入强结构归纳偏置。针对不完美监督和物理不可行性,设计物理感知去噪对齐(PADA),结合可微分有符号距离场(SDF)与测试时优化(TTO),将生成场景投影至物理可行流形,同时保持语义意图。实验表明,PhyScene3D在语义准确性和物理有效性上均优于当前最优方法,场景级碰撞率较人工标注训练数据降低40%。

原文摘要 · Abstract (English)

Generating physically consistent 3D tabletop scenes is a fundamental yet underexplored problem for interactive and generalist robotic learning. The challenge stems from dense object hierarchies and irregular affordances. Here, an interactive scene denotes a physically valid, collision-free environment directly loadable into physics simulators. Existing methods, ranging from decoupled symbolic solvers to end-to-end regression models, often suffer from error propagation or overfitting to noisy supervision containing widespread physical violations. To address these limitations, we introduce PhyScene3D, a framework that reformulates generation as a Human-Mimetic Constructive Process. The proposed Cognitive Topological Reasoning Chain (CTRC) factorizes scene synthesis into a sequential, anchor-conditioned process. It employs a 3D AABB-based placement scheme that imposes a strong structural inductive bias. To address imperfect supervision and physical infeasibility, we introduce Physics-Aware Denoising Alignment (PADA). It integrates a differentiable Signed Distance Field (SDF) with Test-Time Optimization (TTO) to project generated scenes onto a physics-feasible manifold while preserving semantic intent. Experiments demonstrate that PhyScene3D outperforms state-of-the-art approaches in both semantic accuracy and physical validity, achieving a 40% reduction in scene-wise collision rate relative to the human-annotated training data.

3D生成物理一致性机器人仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。