arXiv:2512.08186cs.RO2025-12被引 57

双系统模型让视觉语言导航更智能、更流畅。

Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-and-Language Navigation

  • 分两层:高层推理规划路径,底层快速执行动作
  • 在多个基准上超越以往方法,实测动态环境适应性强
  • 适合做复杂场景的实时导航,如机器人或自动驾驶

尽管近期大视觉语言模型(VLMs)提升了视觉语言导航(VLN)的泛化能力,但现有方法通常依赖端到端流程,将视觉语言输入直接映射为短程离散动作,常导致动作碎片化、延迟高,并难以应对动态障碍物等现实挑战。本文提出首个双系统VLN基础模型DualVLN,协同整合高层推理与低层动作执行。系统2基于VLM的全局规划器,通过图像引导的推理‘慢落地’,预测中期目标点。系统1是轻量级多模态条件扩散变换器策略,利用显式像素目标和系统2的潜在特征,实现平滑精准的轨迹生成。该双系统设计支持复杂动态环境中的鲁棒实时控制与自适应局部决策。通过解耦训练,VLM保持强泛化性,而系统1实现可解释且高效的局部导航。DualVLN在所有VLN基准上均优于先前方法,真实世界实验验证了其在动态环境中具备稳健的长程规划与实时适应能力。

原文摘要 · Abstract (English)

While recent large vision-language models (VLMs) have improved generalization in vision-language navigation (VLN), existing methods typically rely on end-to-end pipelines that map vision-language inputs directly to short-horizon discrete actions. Such designs often produce fragmented motions, incur high latency, and struggle with real-world challenges like dynamic obstacle avoidance. We propose DualVLN, the first dual-system VLN foundation model that synergistically integrates high-level reasoning with low-level action execution. System 2, a VLM-based global planner, "grounds slowly" by predicting mid-term waypoint goals via image-grounded reasoning. System 1, a lightweight, multi-modal conditioning Diffusion Transformer policy, "moves fast" by leveraging both explicit pixel goals and latent features from System 2 to generate smooth and accurate trajectories. The dual-system design enables robust real-time control and adaptive local decision-making in complex, dynamic environments. By decoupling training, the VLM retains its generalization, while System 1 achieves interpretable and effective local navigation. DualVLN outperforms prior methods across all VLN benchmarks and real-world experiments demonstrate robust long-horizon planning and real-time adaptability in dynamic environments.

视觉导航双系统扩散模型机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。