arXiv:2512.19021cs.CVcs.RO2025-12被引 18

构建大规模真实物理模拟的视觉语言导航基准,推动具身智能研究

VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation

  • 统一碎片化任务,支持全运动学真实物理仿真
  • 涵盖大规模多样场景,验证从经典模型到多模态大模型的表现
  • 适合研究具身智能、模拟到现实泛化与通用移动代理的学者

尽管视觉语言导航(VLN)取得显著进展,现有基准仍局限于固定、小规模数据集和简化的物理模拟,限制了对模拟到现实泛化能力的洞察,并造成研究空白。此外,任务分散阻碍了领域内统一进展,而数据规模不足难以满足现代大语言模型预训练需求。为此,我们提出VLNVerse:一个大规模、可扩展的基准,支持多样化、具身化、真实的仿真与评估。其多功能性将此前分散的任务整合至统一框架,并提供可扩展工具包;具身设计采用真实物理引擎,支持全运动学仿真,摆脱传统‘幽灵式’无物理约束的代理。我们利用该基准对现有方法(从经典模型到多模态大模型)进行全面评估,并提出一种新型统一多任务模型,可处理所有任务。VLNVerse旨在缩小模拟导航与真实世界泛化之间的差距,为社区提供关键工具,推动可扩展、通用的具身运动代理研究。

原文摘要 · Abstract (English)

Despite remarkable progress in Vision-Language Navigation (VLN), existing benchmarks remain confined to fixed, small-scale datasets with naive physical simulation. These shortcomings limit the insight that the benchmarks provide into sim-to-real generalization, and create a significant research gap. Furthermore, task fragmentation prevents unified/shared progress in the area, while limited data scales fail to meet the demands of modern LLM-based pretraining. To overcome these limitations, we introduce VLNVerse: a new large-scale, extensible benchmark designed for Versatile, Embodied, Realistic Simulation, and Evaluation. VLNVerse redefines VLN as a scalable, full-stack embodied AI problem. Its Versatile nature unifies previously fragmented tasks into a single framework and provides an extensible toolkit for researchers. Its Embodied design moves beyond intangible and teleporting "ghost" agents that support full-kinematics in a Realistic Simulation powered by a robust physics engine. We leverage the scale and diversity of VLNVerse to conduct a comprehensive Evaluation of existing methods, from classic models to MLLM-based agents. We also propose a novel unified multi-task model capable of addressing all tasks within the benchmark. VLNVerse aims to narrow the gap between simulated navigation and real-world generalization, providing the community with a vital tool to boost research towards scalable, general-purpose embodied locomotion agents.

视觉语言导航具身智能真实仿真多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。