提升英语学习者语音识别公平性,缓解低水平学习者识别效果差的问题
Proficiency-Aware Adaptation and Data Augmentation for Robust L2 ASR
- 根据水平分层设计多任务学习,同时训练语音识别与语言水平分类
- 对低水平语音使用频谱掩码增强,降低错误率最高达29.4%
- 适合教育、无障碍技术中需公平对待非母语用户的场景
通用语音识别在非母语者(如英语二语学习者)上表现不佳,加剧偏见并限制其在教育与无障碍领域的应用。基于CEFR分级的Speak and Improve数据集,我们发现直接微调Whisper虽能降低平均词错误率(WER),却扩大了不同水平学习者间的差距,尤其损害低水平者。为此提出两种策略:(i) 水平感知的多任务学习,联合优化语音识别与水平分类;(ii) 针对性数据增强,对低水平语音施加频谱掩码以缓解数据不平衡。实验显示,该方法使平均WER相对降低29.4%,插入/删除错误相对减少58.6%。关键在于,在反映真实分布的数据严重失衡下,两种策略均持续缩小水平差距,推动面向二语学习者的公平语音识别。
原文摘要 · Abstract (English)
General-purpose ASR underperforms for atypical speakers, such as L2 learners, reinforcing bias and limiting use in education and accessibility. Using the CEFR-graded Speak and Improve corpus, we show that naive fine-tuning of Whisper reduces average WER but simultaneously widens disparities and disproportionately harms lower-level learners. To address this, we propose two strategies: (i) proficiency-aware multitask learning, jointly optimizing ASR with proficiency classification, and (ii) targeted augmentation, applying spectrogram masking to low-proficiency speech to counter imbalance. These approaches reduce WER by up to 29.4 percent (relative) and insertion/deletion errors by as much as 58.6 percent (relative). Crucially, despite the severe imbalance of the dataset reflecting real-world distributions, both strategies consistently narrow proficiency gaps, advancing equitable ASR for L2 learners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。