从单张图生成可模拟的3D资产,关键在于显式物理推理过程。
PhysX-CoT: Structured Physical Reasoning from a Single Image to Simulation-Ready 3D Assets

- 将图像转3D视为分步物理推理链,逐层输出部件状态
- 在多个指标上超越现有基线,几何与物理属性更准确
- 适合需要高精度物理模拟的机器人和具身AI研究
仿真用3D资产是机器人和具身AI的核心。现有方法通常将单图生成视为视觉语言模型输出序列化资产,由解码器生成几何与物理场,但图像到3D的推理过程隐含且不可控。本文提出PhysX-CoT,将单图生成重构为显式的结构化物理推理过程,即按序输出部件级状态轨迹,涵盖分解、2D/3D定位、关系、粗略几何和表面线索,并分别进行监督、条件控制与奖励优化。几何被解耦:3D框负责位置,局部编码负责形状。采用对齐思维链的GRPO优化解析有效性、定位准确性、几何一致性及物理合理性。在统一训练协议下(相同主干网络、数据集、冻结解码器),PhysX-CoT在几何、尺度与物理属性指标上全面优于最接近的全任务基线。对照实验表明显式状态具有功能性而非装饰性;在Unreal Engine 5中,生成资产能正确解析、碰撞与运动。
原文摘要 · Abstract (English)
Simulation-ready 3D assets are central to robotics and embodied AI. Generating them from a single image is usually framed as a vision-language model that emits a serialized asset for a decoder to turn into geometry and physical fields, leaving the image-to-3D reasoning implicit. We argue the limiting factor is this output-centric view: part placement and local shape are entangled in one global-coordinate token stream, and the intermediate physical states are never exposed for supervision, conditioning, or verification. PhysX-CoT instead casts single-image asset generation as an explicit structured physical reasoning process, an ordered and machine-parseable trajectory of part-level states covering decomposition, 2D and 3D grounding, relations, coarse geometry, and surface cues that we separately supervise, use to condition geometry, and treat as reward targets. Geometry is factorized so that 3D boxes carry placement and local codes carry shape, and CoT-aligned GRPO optimizes parse validity, grounding, geometry, placement, and physical consistency. Under a unified protocol that retrains all learned baselines on the same backbone, data, and frozen decoder, PhysX-CoT outperforms the closest full-task baseline across geometry, scale, and physical-attribute metrics. Oracle, token-matched, and state-order controls show the explicit states are functional rather than cosmetic, and in Unreal Engine~5 the generated assets parse, collide, and articulate at high validity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。