arXiv:2410.16712cs.SDcs.CL2024-10中稿 · IEEE ICKG 2024被引 6

通过选择性降噪降低语音识别性别偏差,提升公平性。

DENOASR: Debiasing ASRs through Selective Denoising

  • 基于智能阈值筛选低可懂度语音进行降噪,针对性减少偏差。
  • 在多个数据集上使男女语音识别误差差距平均下降,性能不降。
  • 无需大量人工标注,适合实际部署的公平性优化方案。

自动语音识别(ASR)系统被发现对特定群体存在偏见,受性别、口音和语调等因素影响。噪声对某些口音或语调的说话者影响更大,导致识别误差率不均。本文提出DENOASR框架,一种选择性降噪方法,旨在降低男性与女性之间的词错误率差异。实验表明,结合DEMUCS与LE两种主流降噪技术,可在不牺牲整体性能的前提下有效缓解性别偏差。在TIE、VOX-POPULI、TEDLIUM和FLEURS等多个基准数据集上,使用OpenAI WHISPER和NVIDIA NEMO两款先进开源模型验证,结果表明两性平均词错误率差距显著缩小。降噪仅作用于可懂度低于阈值的语音样本,该阈值由小规模验证集估算得出,避免了大规模人工转录需求。研究证明,选择性降噪是改善当前ASR系统公平性的有效路径。

原文摘要 · Abstract (English)

Automatic Speech Recognition (ASR) systems have been examined and shown to exhibit biases toward particular groups of individuals, influenced by factors such as demographic traits, accents, and speech styles. Noise can disproportionately impact speakers with certain accents, dialects, or speaking styles, leading to biased error rates. In this work, we introduce a novel framework DENOASR, which is a selective denoising technique to reduce the disparity in the word error rates between the two gender groups, male and female. We find that a combination of two popular speech denoising techniques, viz. DEMUCS and LE, can be effectively used to mitigate ASR disparity without compromising their overall performance. Experiments using two state-of-the-art open-source ASRs - OpenAI WHISPER and NVIDIA NEMO - on multiple benchmark datasets, including TIE, VOX-POPULI, TEDLIUM, and FLEURS, show that there is a promising reduction in the average word error rate gap across the two gender groups. For a given dataset, the denoising is selectively applied on speech samples having speech intelligibility below a certain threshold, estimated using a small validation sample, thus ameliorating the need for large-scale human-written ground-truth transcripts. Our findings suggest that selective denoising can be an elegant approach to mitigate biases in present-day ASR systems.

语音识别公平性降噪偏见缓解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。