arXiv:2608.30204cs.CL2026-08中稿 · EMNLP

模型误判讽刺,因依赖夸张语调的错误线索。

When Models Hear What They Expect: Diagnosing Prosodic Heuristics in Multimodal Sarcasm Detection

论文配图:When Models Hear What They Expect: Diagnosing Prosodic Heuristics in Multimodal Sarcasm Detection
图 1 · 摘自论文原文
  • 通过拆解语音与文本贡献,发现模型偏倚语调特征
  • 仅调整语调就使误判率最高升至60%
  • 该现象跨模型存在,揭示通用认知偏差

多模态大模型联合处理语音与文本,但其是否利用语调线索进行语用推断,或仅依赖表面声学模式,尚缺乏系统研究。本文以讽刺检测为任务,评估Qwen2.5-Omni和Qwen3-Omni在中文与英文下的五种模态条件,分解词汇内容、语音语义与语调结构的贡献。加入音频后,误报率系统性上升,但真阳性检测未提升。声学错误诊断显示,模型错误集中于一种共通的表达性语调刻板印象:音高升高与停顿不规则,这与两种语言中真实讽刺线索不符。仅操纵这两个维度即因果验证该启发式,导致误报率高达60%。将相同操控模板应用于Gemini 3 Flash Preview而无需修改,亦重现该效应,表明该刻板印象超越Qwen Omni系列,非单一模型架构所致。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) process speech and text jointly, yet whether they exploit prosodic cues for pragmatic inference or rely on surface acoustic patterns has received little systematic investigation. We address this through sarcasm detection, evaluating Qwen2.5-Omni and Qwen3-Omni on Mandarin Chinese and English under five modality conditions that decompose the contributions of lexical content, vocal semantics, and prosodic structure. Adding audio systematically inflates false positives without improving true positive detection. Acoustic error diagnosis reveals that model errors cluster on a shared stereotype of expressive prosody, namely elevated pitch and irregular pausing, that diverges from the actual cues marking sarcasm in both languages. Targeted manipulation of only these two dimensions causally confirms the heuristic, inducing false positive rates of up to 60%. Applying the same manipulation template to Gemini~3 Flash Preview without modification replicates the effect, suggesting that the stereotype extends beyond the Qwen Omni family rather than arising from a single model architecture.

多模态讽刺识别语音分析模型偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。