arXiv:2508.13603cs.CLcs.AI2025-08被引 4

通过语音角色分配检测语音大模型的性别偏见。

Who Gets the Mic? Investigating Gender Bias in the Speaker Assignment of a Speech-LLM

  • 用语音角色选择作为探测工具,分析模型隐含偏见。
  • 在职业和性别化词汇数据集上测试,发现无系统性偏差。
  • 模型具性别意识但倾向不强,适合研究生成式语音公平性。

与文本类大语言模型类似,语音大模型(Speech-LLMs)也表现出涌现能力与上下文感知。然而,这些特性是否延伸至性别偏见仍不确定。本研究提出一种新方法:利用语音角色分配作为偏见探测工具。不同于文本模型中隐含的性别关联,语音模型需输出具体性别语音,使角色选择成为显式的偏见信号。我们评估了文本转语音模型Bark,分析其对文本提示的默认角色分配。若其角色选择系统性地符合性别刻板印象,则可能反映训练数据或模型设计中的偏见。为此,构建两个数据集:(i) 职业数据集(包含性别刻板职业),(ii) 性别化词汇数据集(带有性别色彩的词语)。结果显示,Bark未表现出系统性性别偏见,但展现出一定的性别敏感性与偏好。

原文摘要 · Abstract (English)

Similar to text-based Large Language Models (LLMs), Speech-LLMs exhibit emergent abilities and context awareness. However, whether these similarities extend to gender bias remains an open question. This study proposes a methodology leveraging speaker assignment as an analytic tool for bias investigation. Unlike text-based models, which encode gendered associations implicitly, Speech-LLMs must produce a gendered voice, making speaker selection an explicit bias cue. We evaluate Bark, a Text-to-Speech (TTS) model, analyzing its default speaker assignments for textual prompts. If Bark's speaker selection systematically aligns with gendered associations, it may reveal patterns in its training data or model design. To test this, we construct two datasets: (i) Professions, containing gender-stereotyped occupations, and (ii) Gender-Colored Words, featuring gendered connotations. While Bark does not exhibit systematic bias, it demonstrates gender awareness and has some gender inclinations.

语音生成性别偏见大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。