通过多组2D-3D映射提升零样本3D定位精度
Multiple Consistent 2D-3D Mappings for Robust Zero-Shot 3D Visual Grounding

- 构建多组一致的2D-3D映射,解决类别错位与几何不准问题
- 在ScanRefer上达62.0% [email protected],领先基线6.4个百分点
- 适合需要高精度3D视觉定位的机器人与具身智能应用
零样本3D视觉定位是开放世界具身AI的关键能力。现有方法受限于开放词汇3D提案质量差,存在类别不准确、几何不精确及多视角推理的空间冗余问题。为此,本文提出MCM-VG框架,通过显式建立多组一致的2D-3D映射,实现鲁棒的零样本3D视觉定位。首先,语义对齐模块利用LLM驱动的查询解析与粗到精的2D-3D匹配纠正类别错位;其次,实例修正模块借助VLM引导的2D分割重建缺失目标,并反投影可靠视觉先验以确立准确3D几何;最后,视角提炼模块聚类3D相机方向,提取最优图像帧。将这些最优RGB帧与鸟瞰图结合成紧凑视觉提示集,将最终目标消歧转化为视觉语言模型的多选推理任务。在ScanRefer和Nr3D基准上的大量实验表明,MCM-VG在零样本3D视觉定位上达到新SOTA。特别地,在ScanRefer上,[email protected]达62.0%,比先前基线高出6.4%;[email protected]达53.6%,领先4.0%。
原文摘要 · Abstract (English)
Zero-shot 3D Visual Grounding (3DVG) is a critical capability for open-world embodied AI. However, existing methods are fundamentally bottlenecked by the poor quality of open-vocabulary 3D proposals, suffering from inaccurate categories and imprecise geometries, as well as the spatial redundancy of exhaustive multi-view reasoning. To address these challenges, we propose MCM-VG, a novel framework that achieves robust zero-shot 3DVG by explicitly establishing Multiple Consistent 2D-3D Mappings. Instead of passively relying on noisy 3D segments, MCM-VG enforces 2D-3D consistency across three fundamental dimensions to achieve precise target localization and reliable reasoning. First, a Semantic Alignment module corrects category mismatches via LLM-driven query parsing and coarse-to-fine 2D-3D matching. Second, an Instance Rectification module leverages VLM-guided 2D segmentations to reconstruct missing targets, back-projecting these reliable visual priors to establish accurate 3D geometries. Finally, to eliminate spatial redundancy, a Viewpoint Distillation module clusters 3D camera directions to extract optimal frames. By pairing these optimal RGB frames with Bird's Eye View maps into concise visual prompt sets, we formulate the final target disambiguation as a multiple-choice reasoning task for Vision-Language Models. Extensive evaluations on ScanRefer and Nr3D benchmarks demonstrate that MCM-VG sets a new state-of-the-art for zero-shot 3D visual grounding. Remarkably, it achieves 62.0\% and 53.6\% in [email protected] and [email protected] on ScanRefer, outperforming previous baselines by substantial margins of 6.4\% and 4.0\%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。