构建物理感知的3D场景图,让机器人理解物体结构与交互关系。
PhysGraph: A Physics-aware 3D Scene Graph for Perception and Reasoning

- 结合符号推理与3D几何,从多视角图像推断物体功能部件和物理属性
- 在合成与真实数据集上实现最佳的语义分割、多物体质量估计和运动关节预测性能
- 适合需要物理合理性与结构化推理的机器人任务,如具身智能与仿真迁移
为完成日常任务,机器人需构建语义丰富、物理可信且结构清晰的3D表示以支持任务规划与可操作性预测。现有方法多关注语义检索,忽视物理与运动学因素;而尝试建模物理特性的方法通常依赖窄域训练集或单物体建模,难以跨物体泛化。为此,我们提出PhysGraph,一种融合符号推理与结构化3D几何的框架,用于在杂乱场景中建模运动学与物理属性。给定RGB-D观测,PhysGraph重建对象中心的3D几何并关联跨视角实例,将物体分解为功能部件,并通过视觉推理推断材料与运动关系。在合成与真实世界数据集上的评估显示,PhysGraph在语义分割、多物体质量估计和关节预测方面达到当前最优表现。其设计简洁高效,生成物理一致且语义结构化的场景图,可用于下游任务,如约束感知的3D可操作性预测和真实到模拟的迁移,实验中已验证其有效性。
原文摘要 · Abstract (English)
To perform a wide range of daily tasks, robots need to construct a 3D representation that is semantically rich, physically grounded, and structured enough to support task planning and affordance prediction. However, existing approaches primarily focus on semantic retrieval, often overlooking physical and kinematic factors. Methods that attempt to model physical properties typically rely on narrow training sets or single-object modeling, limiting scalability and generalization across diverse object types. To address these challenges, we present PhysGraph, a framework that unifies symbolic reasoning with structured 3D geometry to model kinematic and physical properties in cluttered scenes. Given RGB-D observations, PhysGraph reconstructs object-centric 3D geometry and associates object instances across views. It then decomposes objects into functional parts and infers materials and articulations through visual reasoning. Evaluated on both synthetic and real-world datasets, PhysGraph achieves state-of-the-art results in semantic segmentation, multi-object mass estimation, and articulation prediction. With its simple yet effective design, PhysGraph produces physically consistent and semantically structured scene graphs, serving as a structured 3D representation for downstream tasks such as constraint-aware 3D affordance prediction and real-to-sim transfer, both of which are demonstrated in our experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。