用认知框架评估音频模型,发现其能力不均衡。
RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark

- 基于认知心理学构建评估体系,分解听觉智能为五类能力。
- 26个先进模型在不同能力上表现差异显著。
- 适合关注模型真实听觉理解力的研究者使用。
人类通过紧密整合的认知能力(如听觉感知、推理与记忆)处理复杂的听觉环境。尽管大音频语言模型(LALMs)在语音理解与多模态音频推理方面取得进展,现有评估范式仍以任务或模态为中心,仅关注最终性能,忽视底层听觉认知行为。这揭示了人类听觉认知与模型评估之间的根本差距,尤其缺乏将认知原则转化为可操作框架的系统性方法。本文提出RAIL,一种基于卡特尔-霍恩-卡罗尔(CHC)认知框架的人本评估范式。RAIL将听觉认知形式化为五大核心能力,并设计结构化评估任务,探究模型如何处理、保留和整合听觉信息。我们进一步构建了一个认知基础基准,包含严谨的数据筛选与对齐人类评价的协议。对26个顶尖LALMs的评估显示,当前模型在各类认知能力上表现极不均衡。RAIL建立了一种新评估范式,推动评估从任务中心转向认知基础的听觉智能分析。
原文摘要 · Abstract (English)
Humans process rich auditory environments through tightly integrated cognitive capabilities such as audio perception, audio reasoning, and memory. Despite recent progress in large audio-language models (LALMs) across speech understanding and multimodal audio reasoning, current evaluation paradigms remain largely task- or modality-centric, focusing on end performance while overlooking underlying auditory cognitive behaviours. This reveals a fundamental gap between how auditory cognition is understood in humans and how it is evaluated in LALMs, particularly in the lack of frameworks that operationalise cognitive principles beyond task-level metrics to systematically capture model behaviour. In this work, we introduce RAIL, a human-centric evaluation paradigm grounded in the Cattell-Horn-Carroll (CHC) cognitive framework. RAIL formalises auditory cognition into five core capabilities and develop them into structured evaluation tasks that probe how models process, retain, and integrate auditory information. We further construct a cognitively grounded benchmark with principled data curation and human-aligned evaluation protocols. Evaluating 26 state-of-the-art LALMs, we find that current models exhibit highly uneven performance across cognitive abilities. RAIL establishes a new evaluation paradigm that moves beyond task-centric benchmarking toward cognitively grounded assessment of auditory intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。