arXiv:2509.25878cs.CL2025-09被引 1

提升爪哇语和巽他语在噪声下的语音识别鲁棒性

ASR Under Noise: Exploring Robustness for Sundanese and Javanese

  • 用合成噪声和SpecAugment增强训练数据
  • 大模型在噪声下识别准确率显著提升
  • 发现语言特异性问题,适合低资源语言研究者

我们研究了基于Whisper的自动语音识别(ASR)模型在两种主要印度尼西亚地方语言——爪哇语和巽他语中的鲁棒性。尽管近期工作在干净环境下表现出色,但在嘈杂环境下的效果仍不明确。为此,我们实验了多种训练策略,包括合成噪声增强和SpecAugment,并在不同信噪比(SNR)下评估性能。结果表明,噪声感知训练显著提升了鲁棒性,尤其对较大的Whisper模型效果更明显。详细的错误分析揭示了语言特异性挑战,为未来改进指明方向。

原文摘要 · Abstract (English)

We investigate the robustness of Whisper-based automatic speech recognition (ASR) models for two major Indonesian regional languages: Javanese and Sundanese. While recent work has demonstrated strong ASR performance under clean conditions, their effectiveness in noisy environments remains unclear. To address this, we experiment with multiple training strategies, including synthetic noise augmentation and SpecAugment, and evaluate performance across a range of signal-to-noise ratios (SNRs). Our results show that noise-aware training substantially improves robustness, particularly for larger Whisper models. A detailed error analysis further reveals language-specific challenges, highlighting avenues for future improvements

语音识别噪声鲁棒低资源语言Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。