通过模拟真实复杂声学环境,提升语音识别在野外场景的鲁棒性。
Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation

- 构建大规模复合声学数据集,覆盖7类典型噪声与54种组合场景。
- 在多个极端条件测试集上,词错误率降低超30%,优于现有最佳系统。
- 适合需要高鲁棒性语音识别的工业级应用,如智能客服、车载系统。
尽管自动语音识别(ASR)和大型音频语言模型取得快速进展,真实环境中的鲁棒识别仍受限于“声学鲁棒性瓶颈”:模型在严重且复合的声学畸变下常丧失声学依据,导致漏识或幻觉。本文提出Mega-ASR,一个统一的野外环境语音识别框架,结合可扩展的复合数据构建与渐进式声学-语义优化。引入Voices-in-the-Wild-2M数据集,涵盖7类经典声学现象和54种物理上合理的复合场景,并采用声学-语义渐进式监督微调与双粒度词错误率门控策略优化进行训练。大量实验表明,Mega-ASR在恶劣条件下的语音识别基准测试中显著优于现有最先进系统:在VOiCES R4-B-F上词错误率为45.69%(对比54.01%),在NOIZEUS Sta-0上为21.49%(对比29.34%)。在复杂复合声学场景中,相对词错误率降低超过30%,超越强开源与闭源基线,建立了一套可扩展的野外鲁棒语音识别范式。
原文摘要 · Abstract (English)
Despite rapid advances in automatic speech recognition (ASR) and large audio-language models, robust recognition in real-world environments remains limited by an "acoustic robustness bottleneck": models often lose acoustic grounding and produce omissions or hallucinations under severe, compositional distortions. We propose Mega-ASR, a unified ASR-in-the-wild framework that combines scalable compound-data construction with progressive acoustic-to-semantic optimization. We introduce Voices-in-the-Wild-2M, covering 7 classic acoustic phenomena and 54 physically plausible compound scenarios, and train Mega-ASR with Acoustic-to-Semantic Progressive Supervised Fine-Tuning and Dual-Granularity WER-Gated Policy Optimization. Extensive experiments demonstrate that Mega-ASR achieves significant advantages over prior state-of-the-art systems on adverse-condition ASR benchmarks (45.69% vs. 54.01% on VOiCES R4-B-F, and 21.49% vs. 29.34% on NOIZEUS Sta-0). On complex compositional acoustic scenarios, Mega-ASR further delivers over 30% relative WER reduction against strong open- and closed-source baselines, establishing a scalable paradigm for robust ASR in-the-wild.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。