分析主流语音识别系统在性别、口音、年龄上的偏见及碳排放问题
Unveiling Biases while Embracing Sustainability: Assessing the Dual Challenges of Automatic Speech Recognition Systems
- 对比Whisper和MMS在真实场景中的识别偏差
- 发现口音和年龄差异导致识别准确率下降15%以上
- 揭示大模型训练带来的高能耗与碳排放风险
本文聚焦于当前表现优异的自动语音识别(ASR)系统Whisper和Massively Multilingual Speech(MMS),开展偏见与可持续性双重评估。尽管这些系统在受控环境下达到顶尖性能,但其在真实场景下的有效性与公平性仍存疑。研究分析了性别、口音、年龄群体相关的识别偏差,并考察其对下游任务的影响。同时,评估了大型声学模型在训练与推理过程中的碳排放与能源消耗,揭示了高算力需求带来的环境代价。通过实证分析,为当前关于ASR系统偏见与可持续性的讨论提供了重要依据。
原文摘要 · Abstract (English)
In this paper, we present a bias and sustainability focused investigation of Automatic Speech Recognition (ASR) systems, namely Whisper and Massively Multilingual Speech (MMS), which have achieved state-of-the-art (SOTA) performances. Despite their improved performance in controlled settings, there remains a critical gap in understanding their efficacy and equity in real-world scenarios. We analyze ASR biases w.r.t. gender, accent, and age group, as well as their effect on downstream tasks. In addition, we examine the environmental impact of ASR systems, scrutinizing the use of large acoustic models on carbon emission and energy consumption. We also provide insights into our empirical analyses, offering a valuable contribution to the claims surrounding bias and sustainability in ASR systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。