测试大模型是真听懂情绪,还是只靠文字猜。
Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance
- 设计新基准测试,分离文字和声音线索影响。
- 六款模型均严重依赖文字信息,声音线索用得少。
- 适合研究语音情感识别或模型可解释性的读者。
理解语音中的情绪需兼顾文字与声学线索。然而,大型音频语言模型(LALMs)是否真正处理声学信息,仍不明确。本文提出 LISTEN(Narrative 情绪中文字与声学线索的测试),一个受控基准,用于解耦文字依赖与声学敏感性。对六款顶尖 LALMs 的评估显示,模型普遍呈现文字主导:当文字线索中性或缺失时,预测为“中性”;线索一致时性能提升有限;线索冲突时无法准确分类情绪。在非语言语境下,表现接近随机。结果表明当前 LALMs 多数仅“转录”而非“倾听”,过度依赖文字语义而忽视声学线索。LISTEN 为多模态模型的情绪理解评估提供了系统方法。
原文摘要 · Abstract (English)
Understanding emotion from speech requires sensitivity to both lexical and acoustic cues. However, it remains unclear whether large audio language models (LALMs) genuinely process acoustic information or rely primarily on lexical content. We present LISTEN (Lexical vs. Acoustic Speech Test for Emotion in Narratives), a controlled benchmark designed to disentangle lexical reliance from acoustic sensitivity in emotion understanding. Across evaluations of six state-of-the-art LALMs, we observe a consistent lexical dominance. Models predict "neutral" when lexical cues are neutral or absent, show limited gains under cue alignment, and fail to classify distinct emotions under cue conflict. In paralinguistic settings, performance approaches chance. These results indicate that current LALMs largely "transcribe" rather than "listen," relying heavily on lexical semantics while underutilizing acoustic cues. LISTEN offers a principled framework for assessing emotion understanding in multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。