arXiv:2510.18374cs.CL2025-10中稿 · ICASSP 2026被引 1

提升英语语音识别对非母语者的公平性,显著降低不同口音群体的识别误差。

Towards Fair ASR For Second Language Speakers Using Fairness Prompted Finetuning

  • 用轻量适配器融合谱解耦与公平优化策略,改进模型对口音的鲁棒性。
  • 在26个口音组上,宏平均词错误率降低58.7%至58.5%,优于标准微调方法。
  • 适合关注语音识别公平性、多语言应用的研究者与开发者。

本文针对非母语者英语语音识别(ASR)系统的公平性挑战展开研究。对广泛使用的Whisper和Seamless-M4T模型分析发现,26个口音组间词错误率(WER)波动显著,存在明显公平性差距。为此,我们提出基于公平性提示的微调方法,采用轻量级适配器,融合谱解耦(SD)、分组分布鲁棒优化(Group-DRO)与不变风险最小化(IRM)。该方法将传统经验风险最小化(ERM)与交叉熵损失,与公平性驱动目标相结合,在保持整体识别准确率的同时,显著提升各口音组间的公平性。在宏平均词错误率方面,相比预训练的Whisper和Seamless-M4T,分别实现58.7%和58.5%的相对提升;相较于标准交叉熵微调,分别提升9.7%和7.8%。

原文摘要 · Abstract (English)

In this work, we address the challenge of building fair English ASR systems for second-language speakers. Our analysis of widely used ASR models, Whisper and Seamless-M4T, reveals large fluctuations in word error rate (WER) across 26 accent groups, indicating significant fairness gaps. To mitigate this, we propose fairness-prompted finetuning with lightweight adapters, incorporating Spectral Decoupling (SD), Group Distributionally Robust Optimization (Group-DRO), and Invariant Risk Minimization (IRM). Our proposed fusion of traditional empirical risk minimization (ERM) with cross-entropy and fairness-driven objectives (SD, Group DRO, and IRM) enhances fairness across accent groups while maintaining overall recognition accuracy. In terms of macro-averaged word error rate, our approach achieves a relative improvement of 58.7% and 58.5% over the large pretrained Whisper and SeamlessM4T, and 9.7% and 7.8% over them, finetuning with standard empirical risk minimization with cross-entropy loss.

语音识别公平性多口音微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。