arXiv:2608.06718cs.CL2026-08

测试语音模型是否真听懂语气,发现准确率高未必可靠。

Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation

论文配图:Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation
图 1 · 摘自论文原文
  • 用反事实音频干扰测试模型能否识别语气变化
  • 多个模型准确率相似但失败原因不同,隐藏风险
  • 适合评估语音裁判模型的可靠性与安全性

语音-语言模型(ALMs)正被用作语音对话系统评价的裁判,但这类裁判可能并未真正利用语调、情感等副语言信息。本文提出反事实审计方法,固定文本内容,仅改变情感、语调或情感转折时间,迫使裁判关注音频线索而非词汇或表达风格。通过原生单上下文判断协议与对比可恢复性控制评估多种模型(Gemini、GPT及开源音频模型),并分解为感知与响应映射两个能力维度。结果发现:对比成功常高估裁判可靠性;相同整体准确率背后存在不同的失败模式。表明仅靠准确率无法全面评估ALM裁判,部署前需进行深入行为审计。

原文摘要 · Abstract (English)

Audio-language models (ALMs) are increasingly used as judges for speech-to-speech systems, but a judge that receives audio may not actually use paralinguistic evidence. We introduce counterfactual audits for paralinguistic response evaluation. Each audit item holds the transcript fixed while varying affect, prosody, or the timing of an affective shift, forcing a valid judge to track the audio cue rather than lexical content or response style. We evaluate ALM judges using a native one-context judgment protocol and a contrastive recoverability control, then further decompose each item into its constituent perception and response-mapping skills. This yields useful diagnostic states that identify different sources of judge failures. Across Gemini, GPT, and open audio models, we find that contrastive success often overstates native judge reliability, and that similar aggregate accuracies can hide different failure modes. These results suggest that ALM judges should not be evaluated by accuracy alone, instead requiring thorough behavioral audits before deployment.

语音模型副语言模型审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。