提出部件级空间关系建模框架,让3D场景生成更符合物理规律。
PARSE: Part-Aware Relational Spatial Modeling

- 用部件级组装图显式建模物体部件间的几何关系
- 在1万张3D场景数据上训练后,布局推理与部件关系理解显著提升
- 适合做3D场景生成、物理仿真和具身智能的研究者
物体间关系是空间智能的核心,但现有表示方法——如语言介词或对象级场景图——过于粗略,无法精确描述哪些区域实际支撑、包含或接触彼此,导致布局模糊且物理不一致。为解决这一问题,本文提出PARSE框架,显式建模物体部件如何相互作用以确定可行且空间上合理的场景配置。该框架核心是部件中心的组装图(PAG),用于编码特定部件间的几何关系,并通过部件感知的空间配置求解器将这些关系转化为几何约束,生成无碰撞、物理有效的场景。基于此,我们构建了包含10,000个3D室内场景的PARSE-10K数据集,每个场景均来自真实图像布局先验和标注完善的部件形状库,具有密集接触结构及部件级接触图。利用该结构化、空间接地的监督信号,对Qwen3-VL进行微调后,其对象级布局推理能力与部件级关系理解均显著增强;此外,将PAG作为结构先验引入3D生成模型,可生成物理真实度更高、结构更复杂的场景。结果表明,PARSE显著提升了几何接地的空间推理能力,并支持生成物理一致的3D场景。
原文摘要 · Abstract (English)
Inter-object relations underpin spatial intelligence, yet existing representations -- linguistic prepositions or object-level scene graphs -- are too coarse to specify which regions actually support, contain, or contact one another, leading to ambiguous and physically inconsistent layouts. To address these ambiguities, a part-level formulation is needed; therefore, we introduce PARSE, a framework that explicitly models how object parts interact to determine feasible and spatially grounded scene configurations. PARSE centers on the Part-centric Assembly Graph (PAG), which encodes geometric relations between specific object parts, and a Part-Aware Spatial Configuration Solver that converts these relations into geometric constraints to assemble collision-free, physically valid scenes. Using PARSE, we build PARSE-10K, a dataset of 10,000 3D indoor scenes constructed from real-image layout priors and a curated part-annotated shape database, each with dense contact structures and a part-level contact graph. With this structured, spatially grounded supervision, fine-tuning Qwen3-VL on PARSE-10K yields stronger object-level layout reasoning and more accurate part-level relation understanding; furthermore, leveraging PAGs as structural priors in 3D generation models leads to scenes with substantially improved physical realism and structural complexity. Together, these results show that PARSE significantly advances geometry-grounded spatial reasoning and supports the generation of physically consistent 3D scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。