arXiv:2604.17248eess.AScs.CL2026-04被引 3

用真人语音评估大模型生成偏见,更真实反映现实中的刻板印象。

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech

论文配图:VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech
图 1 · 摘自论文原文
  • 基于真人语音和开放任务评估模型偏见
  • 性别与口音引发显著分布偏移,且依赖具体任务
  • 适合关注模型公平性与真实场景测试的研究者

大型音频-语言模型(LALMs)正广泛应用于日常场景,但其生成偏见仍缺乏深入研究。现有语音公平性评测依赖合成语音和多选题(MCQ),难以全面反映公平性问题。我们提出VIBE框架,通过个性化推荐等开放式任务,利用真人录制语音进行评估。相比多选题,该方法允许刻板印象自然显现而无需预设选项,易于拓展至新任务。对12个顶尖LALMs的评估揭示了系统性偏见:性别与口音线索均引发统计显著的分布偏移,且偏见程度强烈依赖任务类型。

原文摘要 · Abstract (English)

Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing speech fairness benchmarks rely on synthetic speech and Multiple-Choice Questions (MCQs), both offering a fragmented view of fairness. We propose VIBE, a framework that evaluates generative bias through open-ended tasks such as personalized recommendations, using human-recorded speech. Unlike MCQs, our method allows stereotypical associations to manifest organically without predefined options, making it easily extensible to new tasks. Evaluating 12 state-of-the-art LALMs reveals systematic biases in realistic scenarios. Both gender and accent cues trigger statistically significant distributional shifts, and bias magnitude is strongly task-dependent.

模型偏见语音评估公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。