构建首个兼顾动态人群与社交距离的导航基准,推动智能体安全避让人类。
HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic Multi-Human Interactions
- 统一离散与连续环境下的社交感知导航任务,定义新评测标准。
- 基于1.68万条带社会情境指令的测试集,发现主流模型在人群干扰下性能骤降。
- 开源数据集、模拟器与排行榜,适合机器人/自动驾驶等场景研究者使用。
视觉语言导航(VLN)长期局限于离散或连续空间,缺乏对动态密集环境的关注。本文提出HA-VLN 2.0,一个统一的基准,引入明确的社交感知约束。贡献包括:(i) 标准化任务与指标,同时评估目标达成率与个人空间遵守情况;(ii) HAPS 2.0数据集与模拟器,建模多主体互动、户外场景及更精细的语言-动作对齐;(iii) 在16,844条具社会语境的指令上进行基准测试,揭示领先智能体在动态人群与部分可观测条件下的性能显著下降;(iv) 通过真实机器人实验验证了从仿真到现实的迁移能力,并开放排行榜实现透明比较。结果表明,显式社交建模能提升导航鲁棒性并减少碰撞,凸显以人为本方法的必要性。通过发布数据集、模拟器、基线与协议,HA-VLN 2.0为安全的人类感知导航研究提供了坚实基础。
原文摘要 · Abstract (English)
Vision-and-Language Navigation (VLN) has been studied mainly in either discrete or continuous spaces, with little attention to dynamic, crowded environments. We present HA-VLN 2.0, a unified benchmark introducing explicit social-awareness constraints. Our contributions are: (i) a standardized task and metrics capturing both goal accuracy and personal-space adherence; (ii) HAPS 2.0 dataset and simulators modeling multi-human interactions, outdoor contexts, and finer language-motion alignment; (iii) benchmarks on 16,844 socially grounded instructions, revealing sharp performance drops of leading agents under human dynamics and partial observability; and (iv) real-world robot experiments validating sim-to-real transfer, with an open leaderboard enabling transparent comparison. Results show that explicit social modeling improves navigation robustness and reduces collisions, underscoring necessity of human-centric approaches. By releasing datasets, simulators, baselines, and protocols, HA-VLN 2.0 provides a strong foundation for safe, human-aware navigation research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。