arXiv:2510.02352cs.CLcs.AI2025-10被引 5

首次系统评估语音对话模型中的偏见,发现开源模型更敏感,推荐任务加剧差异。

Evaluating Bias in Spoken Dialogue LLMs for Real-World Decisions and Recommendations

  • 构建公平性评估框架,用GUS和SNSR量化决策与推荐偏见
  • 闭源模型整体偏见更低,开源模型对年龄性别更敏感
  • 多轮对话中偏见持续存在,推荐任务放大跨群体差异

尽管大语言模型中的偏见(如刻板印象、文化倾向)已被研究,但具有音频输入输出的语音对话模型(SDMs)中的偏见仍缺乏系统探索。声学特征如年龄、性别、口音可能影响模型输出,多轮对话中这些影响可能被放大,进而影响决策与推荐的公平性。本文系统评估了语音LLM中的偏见,使用组不公平得分(GUS)衡量决策偏差,用相似度归一化统计率(SNSR)评估推荐偏差,涵盖Qwen2.5-Omni、GLM-4-Voice等开源模型及GPT-4o Audio、Gemini-2.5-Flash等闭源API。结果表明:闭源模型整体偏见较低,开源模型对年龄和性别更敏感;推荐任务会加剧跨群体差异;多轮对话中偏见仍可持续存在。本研究首次系统分析端到端语音对话模型的偏见,为构建公平可靠的语音交互系统提供依据。为促进后续研究,我们发布了FairDialogue数据集与评估代码。

原文摘要 · Abstract (English)

While biases in large language models (LLMs), such as stereotypes and cultural tendencies in outputs, have been examined and identified, their presence and characteristics in spoken dialogue models (SDMs) with audio input and output remain largely unexplored. Paralinguistic features, such as age, gender, and accent, can affect model outputs; when compounded by multi-turn conversations, these effects may exacerbate biases, with potential implications for fairness in decision-making and recommendation tasks. In this paper, we systematically evaluate biases in speech LLMs and study the impact of multi-turn dialogues with repeated negative feedback. Bias is measured using Group Unfairness Score (GUS) for decisions and similarity-based normalized statistics rate (SNSR) for recommendations, across both open-source models like Qwen2.5-Omni and GLM-4-Voice, as well as closed-source APIs such as GPT-4o Audio and Gemini-2.5-Flash. Our analysis reveals that closed-source models generally exhibit lower bias, while open-source models are more sensitive to age and gender, and recommendation tasks tend to amplify cross-group disparities. We found that biased decisions may persist in multi-turn conversations. This work provides the first systematic study of biases in end-to-end spoken dialogue models, offering insights towards fair and reliable audio-based interactive systems. To facilitate further research, we release the FairDialogue dataset and evaluation code.

语音对话偏见评估公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。