arXiv:2510.25577eess.AScs.AI2025-10中稿 · Interspeech 2026被引 1

提出语音质量评估框架VQ-Bench,揭示语音模型对声调变化的敏感偏差。

Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models

  • 构建合成四类语音特征的平行数据集,可控评估语音模型表现。
  • 发现主流语音模型在声调影响下产生明显人格感知偏差,如领导力与共情度变化。
  • 揭示性别不对称性,适合关注语音AI伦理与公平性的研究者使用。

近期语音基础模型(SFM)可直接处理原始音频,使模型能够响应细微的副语言特征。然而,这些模型如何理解非词汇性线索仍缺乏研究。本文提出VQ-Bench,一个包含合成的清音、气声、破音和末尾破音四种发音类型的控制评估套件。我们在四个生态有效的开放生成任务及语音情绪识别中评估了SFM的敏感性。结果表明存在显著性能差距:领先的商用API未能通过基本生物识别合理性检验,其他模型则在发音类型影响下系统性地改变对代理感、同理心和领导力的判断。研究还揭示了薪资与领导力认可中的性别不对称性,表明语音基础模型可能反映甚至放大人类社会偏见。本工作建立了一个可复现的框架,以确保语音人工智能在副语言理解上的责任性。

原文摘要 · Abstract (English)

Recent advances in Speech Foundation Models (SFMs) enable direct processing of raw audio, allowing models to respond to subtle paralinguistic variation. However, how these models interpret non-lexical cues remains largely unstudied. We introduce VQ-Bench, a controlled evaluation suite featuring a parallel dataset of synthesized modal, breathy, creaky, and end-creak phonation types. We evaluate SFM sensitivity through open-ended generation across four ecologically valid domains, alongside speech emotion recognition. Our results reveal performance gaps: while a leading commercial API failed basic biometric sanity checks, other models exhibited systematic shifts in agency, empathy, and leadership based on phonation. Our findings also highlight gender asymmetries in salary and leadership endorsements, demonstrating that SFMs may mirror or amplify human social biases. This work establishes a reproducible framework for ensuring responsible paralinguistic interpretation in speech-based AI.

语音模型语音质量偏见检测评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。