arXiv:2605.00884cs.CV2026-05

轻量级视觉语言动作模型实现无人机实时导航与语义感知

LiteVLA-H: Dual-Rate Vision-Language-Action Inference for Onboard Aerial Guidance and Semantic Perception

  • 设计双速率系统,快速输出动作指令,慢速生成语义描述
  • 在嵌入式平台实现19.74Hz动作响应与6.08-6.67Hz语义输出
  • 适用于无人机边缘计算场景,兼顾实时性与理解能力

视觉语言动作(VLA)模型在操作任务中展现出强大的语义理解与泛化能力,但其在无人机上的部署仍面临挑战:无人机需在严格限制的机载计算与通信条件下实现低延迟闭环引导。本文提出 LiteVLA-H,一个仅256M参数的紧凑型VLA系统,支持在NVIDIA Jetson AGX Orin平台上双速率运行:快速外环引导模式用于生成短动作标记,较慢的语义模式用于场景理解、风险描述与操作员交互叙述。核心发现是,在此紧凑边缘环境下,端到端延迟主要由多模态预填充阶段决定,而非解码少量额外标记的边际成本。因此,调度器可在50.65ms(19.74Hz)内发出反应式动作标记,同时保持149.90–164.57ms(6.08–6.67Hz)的句级语义输出。为在不牺牲描述能力的前提下优化模型性能,采用一种知识保留微调策略,融合飞行反应数据、航空语义数据及通用图像描述与问答监督。实验表明,该系统在同等部署条件下,动作分支推理速度优于AnywhereVLA、FutureVLA与ReMem-VLA等先进架构,同时维持周期性语义感知能力。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models have shown strong semantic grounding and task generalization in manipulation, but aerial deployment remains difficult because drones require low-latency closed-loop guidance under strict onboard compute and communication constraints. We present LiteVLA-H, a compact 256M-parameter VLA system designed for dual-rate operation on an NVIDIA Jetson AGX Orin: a fast outer-loop guidance mode for short action-token outputs and a slower semantic mode for scene understanding, hazard description, and operator-facing narration. The central empirical observation is that, in this compact edge regime, end-to-end latency is dominated by multimodal pre-fill rather than by the marginal cost of decoding a few extra tokens. This motivates a scheduler that issues reactive action tokens at 50.65,ms (19.74,Hz) while still supporting sentence-level semantic outputs at 149.90--164.57\ms (6.08--6.67,Hz) on the same embedded platform. To specialize the model without collapsing its descriptive competence, we use a knowledge-preserving fine-tuning recipe that mixes reactive flight data, aerial semantic data, and generic caption/VQA supervision. Beyond reporting current latency measurements, we position the system against recent state-of-the-art architectures, including AnywhereVLA, FutureVLA, and ReMem-VLA, showing that the measured action branch reaches a higher edge inference rate under our deployment conditions while retaining periodic semantic awareness.

无人机多模态边缘计算实时推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。