arXiv:2609.02277cs.SDcs.AI2026-09

首个针对大音频语言模型的听觉错觉评测基准,揭示其认知局限。

Auditory Illusion Benchmark for Large Audio Language Models

论文配图:Auditory Illusion Benchmark for Large Audio Language Models
图 1 · 摘自论文原文
  • 构建十类跨音乐、声音、语音的听觉错觉数据集,标注知识先验
  • 模型在低级声学错觉上忠实于信号,高级语义错觉更接近人类
  • 为测试模型认知能力提供新视角,适合研究听觉认知与模型对齐

感知错觉长期以来是探究人类认知的重要工具,揭示了感知中的偏见与局限。在听觉领域,这些错觉为检验大音频语言模型(LALMs)是否复现人类感知倾向提供了独特视角。然而,现有评测多集中于视觉错觉或通用音频任务,听觉错觉仍被严重忽视。为此,我们提出AIB,首个面向大音频语言模型的听觉错觉基准,涵盖音乐、声音和语音中的十类代表性错觉,并为每项标注是否存在基于知识的先验。方法上,将模型评估与受控的人类听觉实验相结合,实现响应的直接对比。结果表明:多数LALMs在低层级声学错觉上保持信号忠实性,但在涉及语言或音乐先验时,部分模型表现出更接近人类的反应,但无一模型能完全匹配人类感知特征。这些发现凸显了当前LALMs作为认知模型的局限性。通过确立听觉错觉作为严谨的测试平台,本工作为探测神经黑箱模型提供了新视角,推动对听觉认知的理解。AIB已公开发布于 https://github.com/gillosae/aib。

原文摘要 · Abstract (English)

Perceptual illusions have long served as crucial probes into human cognition, revealing biases and limitations of perception. In the auditory domain, such illusions provide a unique lens for testing whether Large Audio Language Models (LALMs) replicate human perceptual tendencies. Despite their importance, most benchmarks focus on visual illusions or general audio tasks, leaving auditory illusions underexplored. To this end, we present AIB, the first auditory illusion benchmark for LALMs, covering ten representative illusions across music, sound, and speech, each annotated for the presence of knowledge-based priors. Our methodology pairs model evaluation with controlled human listening studies, enabling direct comparison of responses. Results show systematic differences: while most LALMs remain signal-faithful on low-level acoustic illusions, several exhibit more human-like responses when linguistic or musical priors are involved, although no model matches the human perceptual profile. These findings highlight the current limitations of LALMs as cognitive models. By establishing auditory illusions as a rigorous testbed, our work offers a new perspective for probing neural black-box models and advancing understanding of auditory cognition. AIB is publicly available at https://github.com/gillosae/aib.

听觉错觉大模型评测认知建模音频语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。