让3D模型理解物体部件,提升交互能力
PAR3D: A Unified 3D-MLLM with Part-Aware Representation for Scene Understanding

- 设计部件感知的3D视觉表征,支持细粒度结构建模
- 在新数据集ScenePart上实现部件级问答与指代分割性能提升
- 适合需要精细3D环境理解的机器人、AR/VR应用
近期3D多模态大模型虽实现了统一的3D场景理解,但仍以物体为中心,难以捕捉对具身交互至关重要的细粒度部件结构。本文提出PAR3D,一个统一的部件感知3D-MLLM框架,使模型能够理解、推理并定位3D场景中的物体及其部件。为支持训练与评估,我们构建了ScenePart——一个带部件标注和语言指令的合成3D场景数据集。引入部件感知3D表征学习,增强视觉表示的细粒度语义,并提出分层分割查询生成机制,通过分层对象-部件查询实现部件定位。大量实验表明,该方法显著提升部件级问答与指代分割性能,同时在物体级视觉-语言任务上也表现优异。
原文摘要 · Abstract (English)
Recent advances in 3D multimodal large language models (3D-MLLMs) have enabled unified solutions for 3D scene understanding tasks, including visual question answering, captioning, and referring segmentation. However, existing 3D-MLLMs remain largely object-centric, limiting their ability to model fine-grained part structures that are essential for embodied interaction with 3D environments. In this work, we present PAR3D, a unified part-aware 3D-MLLM framework that enables models to understand, reason about, and ground both objects and their parts in 3D scenes. To enable training and evaluation of part-aware 3D scene understanding, we introduce ScenePart, a synthetic 3D scene dataset with part-level annotations and language instructions. We further develop Part-Aware 3D Representation Learning to enrich 3D visual representations with fine-grained part-level semantics, and propose Hierarchical Segmentation Query Generation to ground part targets via hierarchical object-part queries. Extensive experiments show that our method substantially improves part-level question answering and referring segmentation, while also achieving strong performance across object-level vision-language tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。