让AI能精准定位3D空间并说出具体尺寸,支持对话式测量。
Ground3D-LMM: Fine-Grained 3D Point Grounding and Spatial Reasoning with LMM

- 统一模型结合点云与图像,实现精细3D定位与数值回答。
- 在250万问答数据集上达到91.2%的定位准确率和0.84的测量误差。
- 适合需要空间理解与物理量交互的应用,如机器人导航、智能建筑。
自然语言对3D环境的查询需具备可验证性和度量性才能付诸行动。可验证性要求对目标3D区域进行明确定位,度量性则需以真实单位报告物理尺寸(如大小、厚度、间隙、距离)。现有3D大模型在该领域受限:对话系统通常缺乏显式3D定位,而3D定位模型又不支持交互式、度量感知对话。本文提出Ground3D-LMM,一种统一模型,输入点云及可选RGB图像,支持(i)基于点的响应定位,(ii)在物体及部件级别输出真实单位的度量数值,包括多对象查询。为评估定位与度量的结合,我们定义了3D有根基测量任务,要求预测目标3D区域及对应的真实度量值。我们基于ScanNet和ScanNet++构建大规模数据集,包含密集的对象与部件标注,约250万问答对,涵盖八项任务,并提供人工验证的测试集。多数据集与任务的实验表明,Ground3D-LMM为有根基、度量感知的3D对话理解提供了强基准。数据集与模型已公开。
原文摘要 · Abstract (English)
Natural-language queries about 3D environments become actionable when responses are verifiable and metric. Verifiability requires explicit grounding to the referred 3D region, while metric answers report physical measurements in real-world units (e.g., size, thickness, clearance, and distance). Existing 3D large multimodal models (LMMs) approaches remain limited: conversational systems typically respond without explicit 3D grounding, while 3D grounding models are not designed for interactive, metric-aware dialogue. In this paper, we present Ground3D-LMM, a unified model that takes a point cloud and an optional RGB image as input and supports 3D spatial conversation with (i) point-grounded responses and (ii) metric numeric outputs at both object and part granularity, including multi-object queries. To evaluate this intersection of grounding and measurement, we define the 3D Grounded Measurement task, which requires predicting the referred 3D region and the corresponding metric quantities in real-world units. We introduce a large-scale dataset built on ScanNet and ScanNet++ datasets with dense object and part annotations and roughly 2.5M question-answer pairs spanning eight tasks, along with a manually verified test set. Extensive experiments on multiple datasets and tasks show that our proposed Ground3D-LMM model provides a strong baseline for grounded, metric-aware 3D conversational understanding. Our dataset and model are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。