端到端语音识别让无人机对新手口语指令更懂、更快。
End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users
- 用自监督声学编码+轻量LSTM,不依赖转录直接推理。
- 93%准确率,7毫秒延迟,比传统方法快29倍。
- 跨模态知识蒸馏提升口语鲁棒性,适合真实飞行场景。
语音控制为无人机操作提供了直观替代方案,但现有系统多依赖固定命令词汇,难以处理新手用户的自发、不流畅口语。本文提出一种面向法语实时人机交互的端到端语音理解架构。模型结合冻结的自监督学习声学编码器与轻量级LSTM分类头,并引入跨模态知识蒸馏,将声学表示与文本教师的语义嵌入对齐,无需推理时转录。我们在新构建的VoiceStick语料库上评估,该语料来自29组非专家用户的真实远程操控会话。在简单语音指令上,最佳配置达93%准确率,7毫秒推理延迟,优于级联基线(79%,202毫秒),提速29倍;在完整自发语音测试集上,准确率达82%,跨模态蒸馏始终提升各配置鲁棒性。结果表明,端到端架构在自发语音引导无人机遥操作中不仅可行且更优,兼具语义鲁棒性、低延迟与可校准置信度。
原文摘要 · Abstract (English)
Voice control offers an intuitive alternative to manual drone piloting, yet most existing systems rely on rigid command vocabularies that fail to handle the spontaneous, disfluent speech of naive users. This paper addresses this gap by proposing an End-to-End Spoken Language Understanding architecture for real-time human-drone interaction in French. Our model combines a frozen Self-Supervised Learning acoustic encoder with a lightweight LSTM-based classification head, augmented by a cross-modal knowledge distillation objective that aligns acoustic representations with semantic embeddings from a text teacher, without requiring transcription at inference time. We evaluate our approach on VoiceStick, a novel French corpus of spontaneous speech collected during real teleoperation sessions with 29 nonexpert dyads. On simple voice commands, our best configuration achieves 93% accuracy at 7 ms inference latency, outperforming cascade baselines (79%, 202 ms) with a 29x speedup. On the full spontaneous speech test set, our architecture reaches 82% accuracy, with crossmodal distillation consistently improving robustness across all configurations. These results demonstrate that End-to-End architectures are not only feasible but preferable for spontaneous voice-guided UAV teleoperation, combining semantic robustness, low latency, and calibrated confidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。