arXiv:2606.21340cs.CL2026-06中稿 · Interspeech 2026

用合成语音提升空管语音识别准确率

Synthetic Audio Generation Framework for Air Traffic Control Speech Recognition

论文配图:Synthetic Audio Generation Framework for Air Traffic Control Speech Recognition
图 1 · 摘自论文原文
  • 构建音色模拟管道,生成带口音的空管语音数据
  • 混合真实与合成数据使错误率降低37%
  • 适合需要高精度语音识别的航空领域研究者

尽管自动语音识别系统在通用领域表现优异,但在空管(ATC)场景中仍面临严重挑战,主要由于通道噪声强、非母语(L2)英语口音普遍及数据稀缺。本文提出一种专门针对该问题的合成语音生成框架,通过结合文本转语音、语音转换、L2到L1口音转换以及一种新型可控的L1到L2口音转换技术,模拟真实空管语音的声学特性。在ATCO2数据集上使用Whisper模型进行实验,结果表明:仅使用合成数据微调或混合真实与合成数据微调,均显著优于未微调基线和仅用真实数据微调的基线,词错误率降低37%。

原文摘要 · Abstract (English)

Automatic Speech Recognition (ASR) systems, despite achieving remarkable accuracy in general-purpose domains with native speech (L1), struggle in domains like Air Traffic Control (ATC) due to strong channel noise, a presence of non-native (L2) English accents, and data scarcity. We propose a synthetic data generation pipeline with acoustical properties simulations specifically designed to address this lack of real data to improve recognition accuracy in the ATC domain. Our approach leverages a combination of neural generation techniques, including Text-to-Speech, Voice Conversion, L2-to-L1 accent conversion, and a novel controllable L1-to-L2 accent conversion framework built to simulate accented speech. Our experiments with the Whisper model on the ATCO2 corpus demonstrate that fine-tuning with either synthetic data alone, or a mix of real and synthetic data, significantly improves the word error rate over out-of-the-box and real data only baselines respectively.

语音识别合成数据空管口音转换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。