针对航空管制场景优化自监督语音模型,显著降低识别错误率。
In-domain SSL pre-training and streaming ASR
- 用4500小时空管数据预训练专用语音模型,提升领域适应性。
- 相比通用模型,空管基准测试词错率大幅下降,实测效果更优。
- 支持低延迟流式推理,适合高安全要求的航空应用。
本研究探讨了在航空管制(ATC)环境下,针对特定领域自监督预训练对离线与流式自动语音识别(ASR)的收益。我们在4.5千小时未标注空管语音数据上训练BEST-RQ模型,随后在小规模有标签空管数据集上微调。为实现实时处理,提出采用分块注意力与动态卷积,确保低延迟推理。对比当前最先进的通用语音编码器(如w2v-BERT 2.0和HuBERT),结果表明领域适配的预训练显著提升标准空管基准上的性能,明显降低词错率(WER)。此外,所提出的流式方法在更严格的延迟约束下进一步降低词错率,特别适用于安全关键型航空应用场景。这些发现表明,针对空管数据专门化自监督表示是迈向更准确、高效实际部署语音识别系统的可行路径。
原文摘要 · Abstract (English)
In this study, we investigate the benefits of domain-specific self-supervised pre-training for both offline and streaming ASR in Air Traffic Control (ATC) environments. We train BEST-RQ models on 4.5k hours of unlabeled ATC data, then fine-tune on a smaller supervised ATC set. To enable real-time processing, we propose using chunked attention and dynamic convolutions, ensuring low-latency inference. We compare these in-domain SSL models against state-of-the-art, general-purpose speech encoders such as w2v-BERT 2.0 and HuBERT. Results show that domain-adapted pre-training substantially improves performance on standard ATC benchmarks, significantly reducing word error rates when compared to models trained on broad speech corpora. Furthermore, the proposed streaming approach further improves word error rate under tighter latency constraints, making it particularly suitable for safety-critical aviation applications. These findings highlight that specializing SSL representations for ATC data is a practical path toward more accurate and efficient ASR systems in real-world operational settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。