arXiv:2601.11039cs.SDcs.CL2026-01被引 2

评测大模型对声音物理属性的理解短板,发现其感知能力远低于人类。

SonicBench: Dissecting the Physical Perception Bottleneck in Large Audio Language Models

  • 构建心理物理学基准,分辨音高、响度等12项声音属性。
  • 多数模型表现接近随机,对比任务也未展现人类优势。
  • 音频编码器能捕捉物理信号,瓶颈在后续解码与对齐。

大型音频语言模型(LALMs)在语义和副语言任务上表现优异,但在感知声音基本物理属性(如音高、响度、空间位置)方面仍研究不足。为此,我们提出SonicBench,一个基于心理物理学的基准,系统评估12个核心物理属性在五个感知维度上的表现。与以往数据集不同,SonicBench使用可控生成工具构建刺激材料,支持两种互补范式:识别(绝对判断)与比较(相对判断)。该设计不仅考察感知精度,还测试关系推理能力——人类在此领域通常更具优势。评估结果揭示了LALMs在基础听觉理解上的显著缺陷:多数模型表现接近随机猜测,且反常地未在比较任务中展现预期优势。此外,显式推理带来的提升微乎其微。然而,线性探测分析表明,冻结的音频编码器已成功捕获这些物理线索(准确率至少60%),说明主要瓶颈在于对齐与解码阶段,模型未能有效利用已捕获的感官信号。

原文摘要 · Abstract (English)

Large Audio Language Models (LALMs) excel at semantic and paralinguistic tasks, yet their ability to perceive the fundamental physical attributes of audio such as pitch, loudness, and spatial location remains under-explored. To bridge this gap, we introduce SonicBench, a psychophysically grounded benchmark that systematically evaluates 12 core physical attributes across five perceptual dimensions. Unlike previous datasets, SonicBench uses a controllable generation toolbox to construct stimuli for two complementary paradigms: recognition (absolute judgment) and comparison (relative judgment). This design allows us to probe not only sensory precision but also relational reasoning capabilities, a domain where humans typically exhibit greater proficiency. Our evaluation reveals a substantial deficiency in LALMs' foundational auditory understanding; most models perform near random guessing and, contrary to human patterns, fail to show the expected advantage on comparison tasks. Furthermore, explicit reasoning yields minimal gains. However, our linear probing analysis demonstrates crucially that frozen audio encoders do successfully capture these physical cues (accuracy at least 60%), suggesting that the primary bottleneck lies in the alignment and decoding stages, where models fail to leverage the sensory signals they have already captured.

音频理解感知评测大模型瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。