通过语音转换让用户亲身体验语音偏见,揭示声音特征对AI交互的影响。
From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction

- 用六种口音和两种性别音色构建测试集,评估语音交互中服务质量和内容偏见。
- 发现口音与性别交叉影响模型响应的匹配度和表达长度,存在显著偏差。
- 采用用户互动实验让使用者体验不同声音带来的信任差异,适合研究者参考。
语音大模型直接处理语音输入,但口音和声线特征可能导致偏见行为。现有评估常忽略这些偏见在端到端语音交互中的表现及用户体验。本文区分服务质量差异(如无关或敷衍回复)与连贯输出中的内容偏见,并考察口音与感知性别之间的交叉影响。提出两阶段评估方法:(1) 覆盖六种口音和两种性别呈现的受控测试集,使用无裁判提示-响应指标分析;(2) 采用语音转换技术,让用户以不同声线体验相同内容。两个研究(交互式,N=24;观察式,N=19)显示,语音转换提升对良性回复的信任与可接受性,促进共情;自动分析则揭示了语音大模型在对齐度和冗长性上存在{口音×性别}交叉偏差。结果表明,语音转换可用于探测与体验交叉性语音偏见,本评估体系为语音对话AI提供了更丰富的偏见检测工具。
原文摘要 · Abstract (English)
SpeechLLMs process spoken language directly from audio, but accent and vocal identity cues can lead to biased behaviour. Current bias evaluations often miss how such bias manifests in end-to-end speech interactions and how users experience it. We distinguish quality-of-service disparities (e.g., off-topic or low-effort responses) from content-level bias in coherent outputs, and examine intersectional effects of accent and perceived gender. In this work, we explore a two-part evaluation approach: (1) a controlled test cohort spanning six accents and two gender presentations, analysed with judge-free prompt-response metrics, and (2) an interactive study design using voice conversion to let users experience identical content through different vocal identities. Across two studies (Interactive, N=24; Observational, N=19), we find that voice conversion increases trust and acceptability for benign responses and encourages perspective-taking, while automated analysis in search of quality-of-service disparities, reveals {accent x gender} disparities in alignment and verbosity across SpeechLLMs. These results highlight voice conversion for probing and experiencing intersectional voice bias while our evaluation suite provides richer bias evaluations for spoken conversational AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。