无需训练即可实现通用3D物体定位,解决候选遗漏与证据不全问题。
UniGround: Universal 3D Visual Grounding via Training-Free Scene Parsing
- 基于3D拓扑和多视角语义生成无偏候选对象,不依赖预训练检测器。
- 在ScanRefer上达到46.1% [email protected],ARKitScenes子集达28.7%。
- 零样本适配新场景,对重建噪声和领域偏移具有强鲁棒性。
3D视觉定位(3DVG)从自然语言描述中定位3D场景中的物体,是具身智能应用的基础。尽管基础模型支持开放词汇推理,但通常依赖预生成候选框,产生两个瓶颈:候选瓶颈——特定数据集的3D提案模型在分布外时会遗漏、碎片化或错误分组目标;证据瓶颈——全局渲染保留空间上下文但遮蔽细节,候选中心视图捕捉局部外观却缺乏全局信息。为此,我们提出UniGround,一种零样本3DVG框架,通过全局候选过滤与情境精确定位解决双重瓶颈。全局候选过滤利用3D拓扑与多视角语义线索构建拓扑一致、类别无关的候选,无需数据集训练的3D检测器、任务特异性提案监督或预设框与类别先验。情境精确定位联合推理全局空间上下文与候选中心视觉证据,并通过闭环一致性验证实现可靠目标识别。UniGround在ScanRefer上取得46.1%/34.1% [email protected]/0.5,ARKitScenes子集达28.7% [email protected]。实验表明其在无数据集特异性先验下表现优异,具备跨数据集泛化能力,且对真实重建噪声和实际领域偏移具有鲁棒性。
原文摘要 · Abstract (English)
3D Visual Grounding (3DVG) localizes objects from natural-language descriptions in 3D scenes and is fundamental to embodied AI applications. Although foundation models enable open-vocabulary reasoning, they typically rely on pre-generated candidates, creating two sequential bottlenecks. The \emph{candidate bottleneck} occurs when dataset-specific 3D proposal models miss, fragment, or incorrectly group targets under distribution shifts, excluding them from VLM reasoning. The \emph{evidence bottleneck} stems from incomplete visual evidence: global renderings preserve spatial context but obscure object details, whereas candidate-centric views capture local appearance but lack global context. To address these bottlenecks, we propose UniGround, a zero-shot 3DVG framework that addresses both bottlenecks through Global Candidate Filtering and Contextual Precision Grounding. Global Candidate Filtering constructs topology-consistent, class-agnostic candidates from 3D topology and multi-view semantic cues, without dataset-trained 3D detectors, task-specific proposal supervision, or predefined box and category priors. Contextual Precision Grounding jointly reasons over global spatial context and candidate-centric visual evidence, followed by closed-loop consistency verification for reliable target identification. UniGround achieves 46.1\%/34.1\% [email protected]/0.5 on ScanRefer and 28.7\% [email protected] on the evaluated ARKitScenes subset of EmbodiedScan. Further experiments demonstrate competitive grounding without dataset-specific 3D priors, cross-dataset generalization to unseen indoor scenes, and robustness to real-world reconstruction noise and practical domain shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。