构建可评估视觉语言模型导航能力的高精度基准,支持多类机器人形态。
NaviTrace: Evaluating Embodied Navigation of Vision-Language Models
- 设计新任务:根据指令与机器人类型生成图像空间中的2D导航轨迹。
- 在1000场景、3000条专家轨迹上测试8个先进模型,发现其导航距离人类仍有显著差距。
- 引入语义感知轨迹评分,结合动态时间规整与物体语义惩罚,更贴近人类判断。
视觉语言模型在众多任务和场景中展现出前所未有的性能与泛化能力。将其集成至机器人导航系统,有望推动通用机器人发展。然而,当前评估仍受限于昂贵的真实实验、过于简化的仿真环境以及匮乏的评测基准。本文提出NaviTrace,一个高质量的视觉问答基准:模型接收指令与实体形态(人形、腿式机器人、轮式机器人、自行车),需输出图像空间中的2D导航轨迹。在1000个场景及超过3000条专家轨迹基础上,我们采用新提出的语义感知轨迹评分,综合动态时间规整距离、终点误差与基于像素语义的形态相关惩罚项,该指标与人类偏好高度相关。评估结果揭示了模型普遍存在空间定位与目标识别不足的问题,导致性能持续落后于人类。NaviTrace为真实机器人导航提供了可扩展、可复现的评测标准。基准与排行榜详见 https://leggedrobotics.github.io/navitrace_webpage/。
原文摘要 · Abstract (English)
Vision-language models demonstrate unprecedented performance and generalization across a wide range of tasks and scenarios. Integrating these foundation models into robotic navigation systems opens pathways toward building general-purpose robots. Yet, evaluating these models' navigation capabilities remains constrained by costly real-world trials, overly simplified simulations, and limited benchmarks. We introduce NaviTrace, a high-quality Visual Question Answering benchmark where a model receives an instruction and embodiment type (human, legged robot, wheeled robot, bicycle) and must output a 2D navigation trace in image space. Across 1000 scenarios and more than 3000 expert traces, we systematically evaluate eight state-of-the-art VLMs using a newly introduced semantic-aware trace score. This metric combines Dynamic Time Warping distance, goal endpoint error, and embodiment-conditioned penalties derived from per-pixel semantics and correlates with human preferences. Our evaluation reveals consistent gap to human performance caused by poor spatial grounding and goal localization. NaviTrace establishes a scalable and reproducible benchmark for real-world robotic navigation. The benchmark and leaderboard can be found at https://leggedrobotics.github.io/navitrace_webpage/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。