arXiv:2604.18249cs.CL2026-04被引 1

自监督语音模型从第一层就对特定说话人群体存在偏见,且影响识别与识别任务。

Where Do Self-Supervised Speech Models Become Unfair?

论文配图:Where Do Self-Supervised Speech Models Become Unfair?
图 1 · 摘自论文原文
  • 逐层分析预训练语音模型的公平性,发现偏差从最早隐层就开始出现。
  • 语音识别任务中群体偏差随性能提升而加剧,与识别准确率呈反向关系。
  • 微调无法消除预训练阶段形成的群体偏见,适合关注模型公平性的研究者阅读。

自监督语音编码器模型(S3Ms)对某些说话人群体(SGs)的建模效果优于其他群体,但其技术层面的原因尚不明确。本文首次对预训练自监督语音编码器进行逐层公平性分析,通过在各嵌入层上探测说话人识别(SID)和自动语音识别(ASR)任务的表现。结果发现,所有模型在两个任务中均对某些说话人群体存在嵌入偏差,且这种偏差从最早的隐层即已显现。更关键的是,所有模型在不同任务中表现出相反的层间偏见模式:在最小化整体SID误差的层中,SID偏差被抑制;而在最小化整体ASR误差的层中,ASR偏差却达到最大。这一反向关系在经ASR微调后的模型中依然存在,表明说话人群体层面的偏见是在预训练阶段形成,且难以通过微调消除。

原文摘要 · Abstract (English)

Speech encoder models are known to model members of some speaker groups (SGs) better than others. However, there has been little work in establishing why this occurs on a technological level. To our knowledge, we present the first layerwise fairness analysis of pretrained self-supervised speech encoder models (S3Ms), probing each embedding layer for speaker identification (SID) automatic speech recognition (ASR). We find S3Ms produce embeddings biased against certain SGs for both tasks, starting at the very first latent layers. Furthermore, we find opposite patterns of layerwise bias for SID vs ASR for all models in our study: SID bias is minimized in layers that minimize overall SID error; on the other hand, ASR bias is maximized in layers that minimize overall ASR error. The inverse bias/error relationship for ASR is unaffected when probing S3Ms that are finetuned for ASR, suggesting SG-level bias is established during pretraining and is difficult to remove.

语音模型公平性自监督偏见分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。