arXiv:2603.20743eess.SPcs.SD2026-03被引 1

分析多维社会线索如何共同作用产生语音合成中的性别偏见

The Binding Effect: Analyzing How Multi-Dimensional Cues Form Gender Bias in Instruction TTS

  • 将提示建模为社会地位、职业刻板印象和人格特征的组合
  • 发现多维度交互导致复杂偏见模式,传统单一测试无法捕捉
  • 揭示偏见源于预训练编码器语义先验和数据分布偏差

当前指令式文本转语音(ITTS)中的偏见评估多依赖单变量测试,忽视了社会线索的组合结构。本文将提示建模为社会地位、职业刻板印象与人格描述的组合,分析开源ITTS模型,发现各社会维度间存在系统性交互效应,相互调制形成复杂偏见模式,而单变量基线方法难以识别。关键发现是,这些偏见超越表面表征,与预训练文本编码器的语义先验及训练数据中的偏斜分布密切相关。进一步表明,通用多样性提示无法有效克服这些深层模式,凸显组合式分析对诊断生成语音中潜在风险的必要性。

原文摘要 · Abstract (English)

Current bias evaluations in Instruction Text-to-Speech (ITTS) often rely on univariate testing, overlooking the compositional structure of social cues. In this work, we investigate gender bias by modeling prompts as combinations of Social Status, Career stereotypes, and Persona descriptors. Analyzing open-source ITTS models, we uncover systematic interaction effects where social dimensions modulate one another, creating complex bias patterns missed by univariate baselines. Crucially, our findings indicate that these biases extend beyond surface-level artifacts, demonstrating strong associations with the semantic priors of pre-trained text encoders and the skewed distributions inherent in training data. We further demonstrate that generic diversity prompting is insufficient to override these entrenched patterns, underscoring the need for compositional analysis to diagnose latent risks in generative speech.

语音合成性别偏见多维分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。