针对东南亚口音的空管语音识别,准确率提升至9.82%。
Adapting Automatic Speech Recognition for Accented Air Traffic Control Communications
- 用新数据集微调ASR模型,专攻东南亚口音。
- 在嘈杂环境下实现9.82%的词错误率。
- 为军事等资源受限场景提供实用方案。
空管通信中有效沟通对航空安全至关重要,但现有自动语音识别(ASR)系统仍难以应对东南亚口音带来的挑战,尤其在嘈杂环境中表现不佳。本研究通过构建新数据集,专门微调适用于东南亚口音的ASR模型,显著提升识别准确率,在东南亚口音空管语音上达到9.82%的词错误率(WER)。研究强调区域专属数据集与聚焦口音的训练策略的重要性,为资源有限的军事行动中部署可靠语音识别系统提供了可行路径。结果表明,增强抗噪能力与采用地区化训练数据是改善非西方口音在空管通信中识别准确性的关键。
原文摘要 · Abstract (English)
Effective communication in Air Traffic Control (ATC) is critical to maintaining aviation safety, yet the challenges posed by accented English remain largely unaddressed in Automatic Speech Recognition (ASR) systems. Existing models struggle with transcription accuracy for Southeast Asian-accented (SEA-accented) speech, particularly in noisy ATC environments. This study presents the development of ASR models fine-tuned specifically for Southeast Asian accents using a newly created dataset. Our research achieves significant improvements, achieving a Word Error Rate (WER) of 0.0982 or 9.82% on SEA-accented ATC speech. Additionally, the paper highlights the importance of region-specific datasets and accent-focused training, offering a pathway for deploying ASR systems in resource-constrained military operations. The findings emphasize the need for noise-robust training techniques and region-specific datasets to improve transcription accuracy for non-Western accents in ATC communications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。