arXiv:2604.05830cs.CLcs.AI2026-04中稿 · Speech Language Mo…被引 1

不依赖性别年龄口音信息训练,显著降低语音唤醒的公平性偏差。

"OK Aura, Be Fair With Me": Demographics-Agnostic Training for Bias Mitigation in Wake-up Word Detection

  • 训练时隐藏用户性别、年龄、口音标签,提升模型泛化能力。
  • 最高降低39.94%性别偏差、83.65%年龄偏差、40.48%口音偏差。
  • 适合关注语音系统公平性的研究者与产品开发者。

语音交互广泛应用,但跨不同说话人群体的唤醒词检测仍存在显著人口统计学偏差。本研究评估了无性别、年龄、口音标签的训练方法在缓解此类偏差上的有效性。实验基于OK Aura数据集,采用排除人口属性标签的训练策略,探索数据增强与预训练语音模型知识蒸馏两种技术。结果表明,这些无标签训练方法显著降低偏差:其中一种方法相较基线,性别偏差降低39.94%,年龄偏差降低83.65%,口音偏差降低40.48%。研究证实,标签无关训练可有效提升唤醒词检测的公平性。

原文摘要 · Abstract (English)

Voice-based interfaces are widely used; however, achieving fair Wake-up Word detection across diverse speaker populations remains a critical challenge due to persistent demographic biases. This study evaluates the effectiveness of demographics-agnostic training techniques in mitigating performance disparities among speakers of varying sex, age, and accent. We utilize the OK Aura database for our experiments, employing a training methodology that excludes demographic labels, which are reserved for evaluation purposes. We explore (i) data augmentation techniques to enhance model generalization and (ii) knowledge distillation of pre-trained foundational speech models. The experimental results indicate that these demographics-agnostic training techniques markedly reduce demographic bias, leading to a more equitable performance profile across different speaker groups. Specifically, one of the evaluated techniques achieves a Predictive Disparity reduction of 39.94\% for sex, 83.65\% for age, and 40.48\% for accent when compared to the baseline. This study highlights the effectiveness of label-agnostic methodologies in fostering fairness in Wake-up Word detection.

语音识别公平性唤醒词

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。