语音语言模型对男女声音同问不同答,暴露隐性性别偏见。
Acoustic-based Gender Differentiation in Speech-aware Language Models
- 构建9208条语音样本数据集,分类分析性别相关回应模式。
- 男声倾向响应更明显,而本该区分性别的场景反而忽略性别差异。
- 问题根源在语音编码器生成男性化声学特征,非模型设计缺陷。
语音语言模型(SpeechLMs)通过语音交互革新人机对话,但可能因声学特征导致性别偏差:相同问题因说话人性别不同而获得不同回应。本文提出新数据集,包含9,208个语音样本,分为无性别依赖、性别刻板印象和性别依赖三类。评估LLaMA-Omni系列模型发现悖论现象:整体响应看似一致,实则存在系统性偏差——在性别刻板问题中,所有模型均表现出男性倾向;而在本应依据性别区分的语境中,模型却无视性别。该现象并非源于中性选项或语音性别感知。即使允许中性回答,模型在性别依赖场景仍保持中立。即便对语音进行性别中性化处理,该悖论依然存在。对比带语音编码的SpeechLMs与对应基础大模型,确认偏差主要来自Whisper语音编码器生成的男性化声学标记。研究揭示当前SpeechLM虽强调公平性,却未能合理利用性别信息,亟需更精细的技术来平衡公平与上下文适切性。
原文摘要 · Abstract (English)
Speech-aware Language Models (SpeechLMs) have fundamentally transformed human-AI interaction by enabling voice-based communication, yet they may exhibit acoustic-based gender differentiation where identical questions lead to different responses based on the speaker's gender. This paper propose a new dataset that enables systematic analysis of this phenomenon, containing 9,208 speech samples across three categories: Gender-Independent, Gender-Stereotypical, and Gender-Dependent. We further evaluated LLaMA-Omni series and discovered a paradoxical pattern; while overall responses seems identical regardless of gender, the pattern is far from unbiased responses. Specifically, in Gender-Stereotypical questions, all models consistently exhibited male-oriented responses; meanwhile, in Gender-Dependent questions where gender differentiation would be contextually appropriate, models exhibited responses independent to gender instead. We also confirm that this pattern does not result from neutral options nor perceived gender of a voice. When we allow neutral response, models tends to respond neutrally also in Gender-Dependent questions. The paradoxical pattern yet retains when we applied gender neutralization methods on speech. Through comparison between SpeechLMs with corresponding backbone LLMs, we confirmed that these paradoxical patterns primarily stem from Whisper speech encoders, which generates male-oriented acoustic tokens. These findings reveal that current SpeechLMs may not successfully remove gender biases though they prioritized general fairness principles over contextual appropriateness, highlighting the need for more sophisticated techniques to utilize gender information properly in speech technology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。