提出快慢双系统模型,提升无人机长距离视觉语言导航的稳定性和实时性。
FSD-VLN: Fast-Slow Dual-System Modeling for Aerial Long-Horizon Vision-Language Navigation

- 分设快慢双分支:慢速分支提取语义先验,快速分支生成飞行指令。
- 在未见场景中成功率提升2倍,单步推理延迟降低50%以上。
- 适合需要高稳定性与低延迟的无人机自主导航任务。
视觉语言导航(VLN)通过将语言指令映射到实时视觉输入,实现无人机在未知环境中的自主导航。相比依赖GPS或预编程导航,VLN支持直观的人机交互和更强的环境适应性,需融合高层语义推理与低延迟飞行控制。现有方法在全局多模态理解与序列动作生成间存在结构错位,导致长程空中导航轨迹抖动严重且决策延迟高。为此,我们提出FSD-VLN,一种快-慢双系统架构,解耦语义推理与低延迟飞行指令生成。框架包含两个异步分支:慢速分支从预训练视觉-语言模型中提取稳定的语义先验,快速分支采用扩散变换器(DiT)建模跨时间动作分布,生成一致飞行输出。我们还引入时序感知自适应优化器,以稳定长序列训练并减少梯度振荡。大规模低空仿真实验表明,FSD-VLN在未见场景中导航成功率最高达2倍于当前最优方法,同时单动作推理延迟和总任务运行时间均降低超过50%。本工作验证了解耦语义-控制建模的有效性,为长程空中视觉语言导航提供了实用范式。
原文摘要 · Abstract (English)
Vision-Language Navigation (VLN) enables UAV autonomous navigation in unknown environments by mapping language instructions to real-time visual inputs. Compared with GPS-dependent or pre-programmed navigation, VLN supports intuitive human-machine interaction and stronger environmental adaptability, requiring tight integration of high-level semantic reasoning and low-latency flight control.Existing methods suffer from structural misalignment between global multimodal understanding and sequential action generation, causing jittery trajectories and severe decision latency for long-horizon aerial navigation. To solve this issue, we propose FSD-VLN, a fast-slow dual-system architecture disentangling semantic reasoning and low-latency flight command generation.The framework has two asynchronous branches: a slow stream extracting stable semantic priors from pre-trained vision-language models, and a Diffusion Transformer (DiT) fast stream modeling cross-temporal action distributions to produce consistent flight outputs. We further introduce a time-aware adaptive optimizer to stabilize long-sequence training and reduce gradient oscillation.Large-scale low-altitude simulation experiments show FSD-VLN achieves up to 2X higher navigation success rates on unseen scenes than SOTA methods, while cutting single-action inference delay and total task runtime by over 50%. Our work validates the benefit of decoupled semantic-control modeling and provides a practical paradigm for long-horizon aerial VLN.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。