让无人机在飞行中实时调整控制策略,提升语言指令执行的准确性。
AeroBridge-TTA: Test-Time Adaptive Language-Conditioned Control for UAVs

- 通过语言编码生成目标,结合可在线更新的隐状态实现动态控制。
- 在13种动态差异条件下,平均性能比基线高22.0分,尤其在分布外场景表现突出。
- 适合需要高鲁棒性的无人机自主导航任务,对实际部署有重要价值。
语言引导的无人机常因执行偏差失败:计划轨迹与控制器实际跟踪能力之间存在差距,尤其是在真实动力学与训练时不一致(如质量变化、阻力偏移、执行器延迟、风扰)的情况下。本文提出AeroBridge-TTA,一种语言条件下的测试时自适应控制流程,旨在缩小这一差距。系统包含三部分:语言编码器将指令映射为子目标,一个基于子目标和学习隐状态的自适应策略,以及一个从观测转移数据中在线更新隐状态的测试时适应(TTA)模块。在五项语言控制无人机任务上,面对13种动态差异条件,且使用相同领域随机化设置,AeroBridge-TTA在分布内与强基线PPO-MLP持平,在所有5个分布外(OOD)条件下均胜出,平均高出22.0分(62.7%对比40.7%);整体提升8.5分完全来自分布外表现。相同权重但仅改变步长α的消融实验表明,隐状态更新本身带来4.6倍的分布外性能提升。
原文摘要 · Abstract (English)
Language-guided unmanned aerial vehicles (UAVs) often fail not from bad reasoning or perception, but from execution mismatch: the gap between a planned trajectory and the controller's ability to track it when the real dynamics differ from training (mass changes, drag shifts, actuator delay, wind). We propose AeroBridge-TTA, a language-conditioned control pipeline that targets this gap with test-time adaptation. It has three parts: a language encoder that maps the command into a subgoal, an adaptive policy conditioned on the subgoal and a learned latent, and a test-time adaptation (TTA) module that updates the latent online from observed transitions. On five language-conditioned UAV tasks under 13 mismatch conditions with the same domain randomization, AeroBridge-TTA ties a strong PPO-MLP baseline in-distribution and wins all 5 out-of-distribution (OOD) conditions, +22.0 pts on average (62.7% vs. 40.7%); the +8.5 pt overall gain comes entirely from the OOD regime. A same-weights ablation that only changes the step size $α$ shows the latent update itself is responsible for a $4.6\times$ OOD lift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。