首个统一多任务的视频式导航模型,支持真实场景下复杂指令无缝执行。
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

- 统一多种导航任务的输入输出格式,实现端到端联合建模。
- 在4个子任务上收集360万数据样本,实现在未见环境中的强泛化能力。
- 适合需要跨任务、长时序、真实世界部署的智能体研究与应用。
实际导航智能体需应对多样交互需求,如遵循指令、搜寻物品、回答问题、追踪人员等。现有具身导航模型难以作为实用通用模型,常受限于特定任务配置或预设地图的离散路径点。本文提出Uni-NaVid,首个基于视频的视觉-语言-动作(VLA)模型,旨在统一多种具身导航任务,实现对未见真实环境中的混合长时序任务无缝导航。该模型通过统一各类任务的输入输出配置,将所有任务整合于单一模型中。训练过程中,从四个核心导航子任务共收集360万条导航数据样本,并促进任务间的协同学习。在综合性导航基准上的大量实验清晰展示了统一建模的优势,性能达到当前最优。此外,真实世界实验验证了模型的有效性与高效性,凸显其强大泛化能力。
原文摘要 · Abstract (English)
A practical navigation agent must be capable of handling a wide range of interaction demands, such as following instructions, searching objects, answering questions, tracking people, and more. Existing models for embodied navigation fall short of serving as practical generalists in the real world, as they are often constrained by specific task configurations or pre-defined maps with discretized waypoints. In this work, we present Uni-NaVid, the first video-based vision-language-action (VLA) model designed to unify diverse embodied navigation tasks and enable seamless navigation for mixed long-horizon tasks in unseen real-world environments. Uni-NaVid achieves this by harmonizing the input and output data configurations for all commonly used embodied navigation tasks and thereby integrating all tasks in one model. For training Uni-NaVid, we collect 3.6 million navigation data samples in total from four essential navigation sub-tasks and foster synergy in learning across them. Extensive experiments on comprehensive navigation benchmarks clearly demonstrate the advantages of unification modeling in Uni-NaVid and show it achieves state-of-the-art performance. Additionally, real-world experiments confirm the model's effectiveness and efficiency, shedding light on its strong generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。