arXiv:2510.01254cs.CLcs.AI2025-10中稿 · IEEE ICASSP 2026被引 6

语音大模型的偏见测试基准难以推广到其他任务。

Do Bias Benchmarks Generalise? Evidence from Voice-based Evaluation of Gender Bias in SpeechLLMs

  • 用LoRA微调模型诱导特定回答倾向。
  • 跨任务测试发现基准表现不一致,长文本生成更差。
  • 提出新评估套件以测模型行为迁移能力。

近期语音大语言模型(SpeechLLMs)的偏见与公平性评测主要依赖多项选择题问答(MCQA)格式:模型需在刻板印象、反刻板印象或中立/无关答案间选择。此类基准隐含假设——模型在不同MCQA任务、语音风格及长篇生成等真实场景下表现一致。本文通过LoRA适配器微调三款SpeechLLMs,使其分别偏好刻板、反刻板或中立答案,并验证这些行为是否泛化至另一独立MCQA基准和长篇创造性生成任务。结果表明,现有MCQA基准无法可靠预测模型在其他MCQA或长篇任务中的表现。结论指出当前语音领域MCQA偏见基准缺乏跨任务泛化能力,并提出一套评估体系以衡量未来模型与评测框架的行为可迁移性。

原文摘要 · Abstract (English)

Recent work in benchmarking bias and fairness in speech large language models (SpeechLLMs) has relied heavily on multiple-choice question answering (MCQA) formats. The model is tasked to choose between stereotypical, anti-stereotypical, or neutral/irrelevant answers given an input speech prompt and an optional text prompt. Such MCQA benchmarks implicitly assume that model performance is consistent across other MCQA tasks, voices, and other task formats such as more realistic, long-form evaluations. In this paper, we probe that assumption. We fine-tune three SpeechLLMs using LoRA adapters to induce specific MCQA behaviours: preference for stereotypical, anti-stereotypical, or neutral/uncertain answers. We then evaluate whether these behaviours generalise to another, distinct MCQA benchmark, and more critically to long-form, creative generation tasks. Our results show that performance on MCQA bias benchmarks fails to reliably predict performances across other MCQA benchmarks, and more importantly across long-form tasks. We conclude that current MCQA bias benchmarks show limited evidence of cross-task generalisation in the speech domain, and also propose an evaluation suite for measuring behaviour transferability in future models and benchmarks.

语音模型偏见评测行为迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。