arXiv:2604.24401cs.SDcs.AI2026-04被引 2

发现大模型答题常靠文本而非听觉,评测可能误判理解能力。

All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation

论文配图:All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation
图 1 · 摘自论文原文
  • 用文本先验和音频依赖双维度诊断模型真实听觉理解力。
  • 无音频时模型仍保持60%-72%原音频得分,说明严重依赖文本。
  • 超95%需音频的问题只需片段即可解决,证明评测设计存漏洞。

大型音频-语言模型在语音与音频基准测试中表现持续提升,但高分未必反映真实的听觉理解能力。若模型无需处理声学信号即可作答,该评测便无法有效衡量听觉理解。本文提出一种诊断框架,包含两个维度:文本先验(仅凭文本和通用知识判断回答能力)和音频依赖性(评估对声学信号的实际依赖程度)。在三个基准上评估八种大型音频-语言模型,发现即使没有音频输入,模型仍能保留60%-72%的完整音频得分。此外,在需要音频的任务中,仅有3.0%-4.2%需完整音频片段,多数问题可通过局部音频片段解决。这些结果挑战了‘性能优异即代表强音频理解’的假设,并提出改进评测可靠性和基准设计的实践建议。

原文摘要 · Abstract (English)

Large Audio-Language Models show consistent performance gains across speech and audio benchmarks, yet high scores may not reflect true auditory perception. If a model can answer questions without processing the acoustic signal, the benchmark fails as a measure of auditory understanding. We present a diagnostic framework using two axes: text prior, which measures answerability from text and general knowledge alone, and audio reliance, which assesses actual dependency on the acoustic signal. Evaluating eight LALMs across three benchmarks, we find that models retain 60-72% of their full audio scores even without any audio input. Moreover, among items that require audio, only 3.0-4.2% need the complete audio clip; the majority can be resolved using localized fragments. These findings challenge the assumption that benchmark performance equals robust audio understanding, and we conclude with practical guidelines for improving evaluation reliability and benchmark design.

音频理解评测诊断文本先验模型偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。