让无人机像侦察蜂一样主动找证据,精准回答开放世界问题。
ScoutVLA: UAV-Centric Active Perception via a Dual-Expert VLA Model for Open-World Embodied Question Answering

- 分双专家结构:视觉语言专家识需求,动作专家生成连续视角调整轨迹。
- 实测成功率达基线10.48倍,问答正确率提升7.72倍。
- 适合需要精细环境感知的无人机智能任务,如搜救与巡检。
空中具身问答(EQA)要求无人机主动感知环境并回答自然语言问题。现有系统通常在目标进入视野后即停止,难以应对需细粒度视角调整的证据探寻类问题。为此,我们提出FG-EQA基准,包含超过4万条仿真轨迹和1000条真实轨迹。受侦察蜂‘摇摆舞’启发,我们设计ScoutVLA——一种以证据驱动的视觉-语言-动作模型。其采用解耦双专家架构:视觉语言专家识别缺失证据的语义意图,独立动作专家利用高自由度流匹配生成连续视角优化轨迹。为平衡连续控制与语义推理,提出知识隔离训练策略,防止动作梯度侵蚀多模态推理能力。大量仿真实验与真实场景测试表明,ScoutVLA相比最先进基线平均严格成功率提升10.48倍,平均问答正确率提升7.72倍。
原文摘要 · Abstract (English)
Aerial Embodied Question Answering (EQA) requires Unmanned Aerial Vehicles (UAVs) to actively perceive the environment and answer natural language questions. Existing outdoor EQA systems usually stop once the target enters the UAV's field of view, leaving the fine-grained viewpoint adjustment needed for evidence-seeking questions largely unresolved. To address this issue, we introduce FG-EQA, a fine-grained active perception EQA benchmark with more than 40K simulated trajectories and 1K real-world trajectories. Drawing inspiration from the ``waggle dance'' of scout bees, which iteratively adjust their flight paths to verify target information, we propose ScoutVLA, an evidence-driven Vision-Language-Action model for outdoor EQA. To emulate this active exploration behavior, ScoutVLA features a decoupled dual-expert architecture: a vision-language expert infers the semantic intent to identify missing evidence, while an independent action expert employs high-DoF flow matching to generate continuous viewpoint-refinement trajectories. To balance the competing demands of continuous control and semantic reasoning, we devise a decoupled training strategy with a knowledge insulation mechanism that prevents the action gradients from erasing the model's multimodal reasoning ability. Extensive simulated experiments and a qualitative real-world field study both verify the superiority of ScoutVLA over the state-of-the-art baselines, demonstrating a 10.48$\boldsymbol{\times}$ higher average strict success rate and a 7.72$\boldsymbol{\times}$ higher average QA correctness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。