arXiv:2603.16941eess.AScs.CL2026-03中稿 · Interspeech 2026被引 1

通过语音克隆控制语言内容,量化评估三款语音大模型的音色与性别交叉偏见。

The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs

  • 用语音克隆保持语义不变,测试六种英语口音与两种性别表达下的模型响应差异。
  • 东欧口音(尤其女性)获得显著更低的帮助度评分,但回应仍保持礼貌。
  • 人类评估者比模型评委更敏锐,能捕捉更强的口音差异,揭示隐藏偏见。

语音大语言模型(SpeechLLMs)直接处理语音输入,保留了口音和感知性别等特征,这些特征在以往级联流程中已被去除。这导致响应结果随说话人身份产生差异。本文基于2,880次受控交互,对三种SpeechLLMs进行大规模交叉性评估,涵盖六种英语口音与两种性别呈现方式,通过语音克隆保持语言内容恒定。采用点态评分、成对比较及最佳最差选择法,并经人工验证,检测出持续存在的方向性偏差:东欧口音的语音获得较低帮助度评分,尤其在女性呈现时更为明显。尽管模型评标捕捉到了偏见的方向趋势,人类评估者表现出显著更高的敏感度,凸显更强的口音层级对比。

原文摘要 · Abstract (English)

Speech Large Language Models (SpeechLLMs) process spoken input directly, retaining cues such as accent and perceived gender that were previously removed in cascaded pipelines. This introduces speaker identity dependent variation in responses. We present a large-scale intersectional evaluation of accent and gender bias in three SpeechLLMs using 2,880 controlled interactions across six English accents and two gender presentations, keeping linguistic content constant through voice cloning. Using pointwise LLM-judge ratings, pairwise comparisons, and Best-Worst Scaling with human validation, we detect recurring directional disparities. Eastern European-accented speech receives lower helpfulness scores, particularly for female-presenting voices. Responses remain polite but differ in helpfulness. While LLM judges capture the directional trend of these biases, human evaluators exhibit significantly higher sensitivity, showing stronger accent-level contrasts.

语音大模型交叉偏见口音歧视评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。