arXiv:2509.14063cs.RO2025-09中稿 · the 2026 IEEE Inte…被引 1

用飞行员对话提升无人机目标预测准确率

Language Conditioning Improves Accuracy of Aircraft Goal Prediction in Non-Towered Airspace

  • 结合语音识别与语言模型解析飞行员通话,提取飞行意图
  • 融合语言信息后,目标预测误差显著低于仅靠轨迹的方法
  • 适合需要理解人类交流的自主飞行系统研发者

自主飞行器需在无塔台管制空域中安全运行,此类环境依赖飞行员间语音沟通协调。为保障安全,飞行器须预测其他飞机的飞行意图及目标位置。本文提出一种多模态框架,将自然语言理解与空间推理结合,提升自主决策能力。通过自动语音识别与大语言模型对飞行员无线电通话进行转录与语义解析,识别飞机身份并提取离散意图标签。这些意图标签与观测轨迹融合,用于条件化时序卷积网络与高斯混合模型,实现概率化目标预测。实验表明,相较仅依赖运动历史的基线方法,本方法显著降低目标预测误差。基于真实非塔台机场数据集的测试验证了该方法的有效性,展示了其在实现社会感知、语言驱动的机器人运动规划中的潜力。

原文摘要 · Abstract (English)

Autonomous aircraft must safely operate in non-towered airspace, where coordination relies on voice-based communication among human pilots. Safe operation requires an aircraft to predict the intent, and corresponding goal location, of other aircraft. This paper introduces a multimodal framework for aircraft goal prediction that integrates natural language understanding with spatial reasoning to improve autonomous decision-making in such environments. We leverage automatic speech recognition and large language models to transcribe and interpret pilot radio calls, identify aircraft, and extract discrete intent labels. These intent labels are fused with observed trajectories to condition a temporal convolutional network and Gaussian mixture model for probabilistic goal prediction. Our method significantly reduces goal prediction error compared to baselines that rely solely on motion history, demonstrating that language-conditioned prediction increases prediction accuracy. Experiments on a real-world dataset from a non-towered airport validate the approach and highlight its potential to enable socially aware, language-conditioned robotic motion planning.

飞行预测多模态语言理解自主飞行

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。