首次发现语音大模型在性别偏见测试中存在声音性别差异的定位偏差。
When Voice Matters: Evidence of Gender Disparity in Positional Bias of SpeechLLMs
- 通过语音输入和提示设计分析模型温度对性别与位置偏见的影响。
- 女性语音引发的偏见强度显著高于男性语音,且位置偏见在语音领域同样存在。
- 当前多选题评测无法捕捉语音偏见,需改进评估方法以保障公平性。
基于语音大模型的对话系统快速发展,亟需可靠基准测试以评估其公平性与偏见。现有基准多依赖多选题问答(MCQA),本文首次开展基于分词概率与响应行为的分析,研究三个问题:1)模型温度与提示设计如何影响MCQA中的性别与位置偏见;2)输入语音性别如何影响这些偏见;3)观察到的趋势是否在另一性别偏见基准中复现。结果表明,文本领域的位置偏见在语音领域同样显著,且对女性语音的影响更强烈。本研究首次在语音大模型性别偏见基准中分离出位置偏见效应,证实当前MCQA基准未考虑语音特性,亟需新评估策略以确保所有用户公平性。
原文摘要 · Abstract (English)
The rapid development of SpeechLLM-based conversational AI systems has created a need for robustly benchmarking these efforts, including aspects of fairness and bias. At present, such benchmarks typically rely on multiple choice question answering (MCQA). In this paper, we present the first token-level probabilistic evaluation and response-based study of several issues affecting the use of MCQA in SpeechLLM benchmarking: 1) we examine how model temperature and prompt design affect gender and positional bias on an MCQA gender-bias benchmark; 2) we examine how these biases are affected by the gender of the input voice; and 3) we study to what extent observed trends carry over to a second gender-bias benchmark. Our results show that concerns about positional bias from the text domain are equally valid in the speech domain. We also find the effect to be stronger for female voices than for male voices. To our knowledge, this is the first study to isolate positional bias effects in SpeechLLM-based gender-bias benchmarks. We conclude that current MCQA benchmarks do not account for speech-based bias and alternative strategies are needed to ensure fairness towards all users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。