arXiv:2511.01755cs.CVcs.RO2025-11NeurIPS被引 11

构建了超大规模3D视觉定位数据集,支持车、无人机、四足机器人多平台跨域学习。

3EED: Ground Everything Everywhere in 3D

  • 融合多平台RGB与激光雷达数据,构建跨场景3D语言定位基准
  • 包含12.8万物体和2.2万有效指代表达,规模达现有数据集10倍
  • 提出平台感知归一化与跨模态对齐技术,助力跨设备泛化研究

3D视觉定位是智能体在开放世界环境中定位语言指代物体的关键。然而,现有基准存在局限于室内场景、单一平台和小规模等问题。本文提出3EED,一个跨平台、多模态的3D定位基准,涵盖来自车辆、无人机和四足机器人平台的RGB与激光雷达数据。该数据集包含超过128,000个物体和22,000个经验证的指代表达,覆盖多样化的室外场景,规模为现有数据集的10倍。我们设计了一种可扩展的标注流程,结合视觉-语言模型提示与人工验证,确保空间定位的高质量。为支持跨平台学习,提出平台感知归一化与跨模态对齐方法,并建立域内与跨平台评估协议。实验发现显著性能差距,揭示通用3D定位的挑战与机遇。3EED数据集与基准工具包已开源,以推动语言驱动的3D具身感知研究。

原文摘要 · Abstract (English)

Visual grounding in 3D is the key for embodied agents to localize language-referred objects in open-world environments. However, existing benchmarks are limited to indoor focus, single-platform constraints, and small scale. We introduce 3EED, a multi-platform, multi-modal 3D grounding benchmark featuring RGB and LiDAR data from vehicle, drone, and quadruped platforms. We provide over 128,000 objects and 22,000 validated referring expressions across diverse outdoor scenes -- 10x larger than existing datasets. We develop a scalable annotation pipeline combining vision-language model prompting with human verification to ensure high-quality spatial grounding. To support cross-platform learning, we propose platform-aware normalization and cross-modal alignment techniques, and establish benchmark protocols for in-domain and cross-platform evaluations. Our findings reveal significant performance gaps, highlighting the challenges and opportunities of generalizable 3D grounding. The 3EED dataset and benchmark toolkit are released to advance future research in language-driven 3D embodied perception.

3D定位多模态具身智能跨平台

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。