发现语音质量评分中存在性别偏差,提出针对性改进模型
MOS-Bias: From Hidden Gender Bias to Gender-Aware Speech Quality Assessment
- 通过分析人类评分发现男性比女性更倾向给高分
- 低质量语音下男女评分差距最大,随质量提升逐渐缩小
- 提出基于群体嵌入的性别感知模型,提升公平性与准确率
平均意见分(MOS)是语音质量评估的标准指标,但人类标注中的偏见尚未被充分研究。我们首次系统分析了MOS中的性别偏见,发现男性听者始终比女性听者给出更高评分——这一差距在低质量语音中最为显著,并随语音质量提升逐渐缩小。这种质量依赖性的结构难以通过简单校准消除。我们进一步证明,基于聚合标签训练的自动MOS模型其预测结果偏向男性感知标准。为此,我们提出一种性别感知模型,通过抽象二元群体嵌入学习性别特异性评分模式,从而提升整体及性别细分下的预测准确性。本研究表明,MOS中的性别偏见是一种系统性、可学习的模式,亟需在公平语音评估中予以重视。
原文摘要 · Abstract (English)
The Mean Opinion Score (MOS) serves as the standard metric for speech quality assessment, yet biases in human annotations remain underexplored. We conduct the first systematic analysis of gender bias in MOS, revealing that male listeners consistently assign higher scores than female listeners--a gap that is most pronounced in low-quality speech and gradually diminishes as quality improves. This quality-dependent structure proves difficult to eliminate through simple calibration. We further demonstrate that automated MOS models trained on aggregated labels exhibit predictions skewed toward male standards of perception. To address this, we propose a gender-aware model that learns gender-specific scoring patterns through abstracting binary group embeddings, thereby improving overall and gender-specific prediction accuracy. This study establishes that gender bias in MOS constitutes a systematic, learnable pattern demanding attention in equitable speech evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。