构建语音模型偏见测试集,区分内容与声音带来的刻板印象。
VoiceBBQ: Investigating Effect of Content and Acoustics in Social Bias of Spoken Language Model
- 将文本偏见测试扩展为可控语音条件,分离内容与声学影响。
- 发现不同模型对性别、口音偏见的响应差异显著,有的放大有的抑制。
- 适合研究语音模型公平性或评估多模态系统偏见的学者使用。
我们提出 VoiceBBQ,即 BBQ(问答偏见基准)的语音扩展版本,通过提供含歧义或消歧语境的语音样本,检测语音语言模型中的社会偏见。由于语音特性,模型偏见可能源于内容和声学两个方面。该数据集将每个 BBQ 语境转换为受控语音条件,支持对内容与声学维度的准确率、偏见度和一致性评分的独立分析,且结果可与原文本基准直接比较。利用 VoiceBBQ,我们评估了两个语音语言模型:LLaMA-Omni 在声学上表现鲁棒,但加剧了性别和口音偏见;而 Qwen2-Audio 显著减弱了这些声学线索,同时保持内容准确性。因此,VoiceBBQ 提供了一个紧凑、可直接插入的测试平台,用于联合诊断语音语言模型在内容与声学层面的偏见。
原文摘要 · Abstract (English)
We introduce VoiceBBQ, a spoken extension of the BBQ (Bias Benchmark for Question Answering) - a dataset that measures social bias by presenting ambiguous or disambiguated contexts followed by questions that may elicit stereotypical responses. Due to the nature of speech, social bias in Spoken Language Models (SLMs) can emerge from two distinct sources: 1) content aspect and 2) acoustic aspect. The dataset converts every BBQ context into controlled voice conditions, enabling per-axis accuracy, bias, and consistency scores that remain comparable to the original text benchmark. Using VoiceBBQ, we evaluate two SLMs - LLaMA-Omni and Qwen2-Audio - and observe architectural contrasts: LLaMA-Omni resists acoustic bias while amplifying gender and accent bias, whereas Qwen2-Audio substantially dampens these cues while preserving content fidelity. VoiceBBQ thus provides a compact, drop-in testbed for jointly diagnosing content and acoustic bias across spoken language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。