研究指令式语音合成模型对职业描述的性别偏见,发现模型易强化刻板印象。
Gender Bias in Instruction-Guided Speech Synthesis Models
- 用职业提示词测试模型生成语音的性别倾向
- 部分职业提示下模型显著偏向特定性别表达
- 不同规模模型偏见程度差异明显,适合评估伦理风险
近年来,可控表达式语音合成技术(尤其是文本到语音,TTS)的发展使得通过文本描述生成特定风格语音成为可能,即风格提示。尽管这提升了语音合成的灵活性与自然度,但人们对模型如何处理模糊或抽象风格提示的理解仍不充分。本研究探讨了模型在解读与职业相关的提示时是否存在性别偏见,重点考察其对“像护士一样说话”等指令的响应。实验结果表明,模型在某些职业提示下存在明显的性别偏向;此外,不同规模的模型在这些职业上的偏见程度各不相同。
原文摘要 · Abstract (English)
Recent advancements in controllable expressive speech synthesis, especially in text-to-speech (TTS) models, have allowed for the generation of speech with specific styles guided by textual descriptions, known as style prompts. While this development enhances the flexibility and naturalness of synthesized speech, there remains a significant gap in understanding how these models handle vague or abstract style prompts. This study investigates the potential gender bias in how models interpret occupation-related prompts, specifically examining their responses to instructions like "Act like a nurse". We explore whether these models exhibit tendencies to amplify gender stereotypes when interpreting such prompts. Our experimental results reveal the model's tendency to exhibit gender bias for certain occupations. Moreover, models of different sizes show varying degrees of this bias across these occupations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。