arXiv:2601.16540cs.SDcs.AI2026-01

探究音频大模型与人脑听觉神经活动的对齐程度

Do Models Hear Like Us? Probing the Representational Alignment of Audio LLMs and Naturalistic EEG

  • 用8种相似性度量对比12个开源音频大模型与脑电数据
  • 发现模型在250-500毫秒间与人脑神经活动高度对齐
  • 负面语调会降低几何相似性但增强相关性,揭示情感处理差异

音频大语言模型(Audio LLMs)在融合语音感知与语言理解方面表现出色,但其内部表征在自然聆听过程中是否与人类神经动态对齐仍不清楚。本文系统考察了12个开源Audio LLMs与脑电图(EEG)信号在两个数据集上的逐层表征对齐情况,采用8种相似性度量(如基于斯皮尔曼的相关性分析,RSA)刻画句内表征几何结构。研究发现三个关键结果:(1) 模型排名在不同度量下存在显著差异,呈现依赖排序特性;(2) 观察到时空对齐模式,表现为深度相关的对齐峰值,并在250-500毫秒时间窗内出现明显的RSA上升,符合与N400相关的神经动力学特征;(3) 发现情感分离现象:通过提出的三模态邻域一致性(TNC)标准识别出负向语调后,几何相似性下降,而基于协方差的依赖关系增强。这些发现为Audio LLMs的表征机制提供了新的神经生物学见解。

原文摘要 · Abstract (English)

Audio Large Language Models (Audio LLMs) have demonstrated strong capabilities in integrating speech perception with language understanding. However, whether their internal representations align with human neural dynamics during naturalistic listening remains largely unexplored. In this work, we systematically examine layer-wise representational alignment between 12 open-source Audio LLMs and Electroencephalogram (EEG) signals across 2 datasets. Specifically, we employ 8 similarity metrics, such as Spearman-based Representational Similarity Analysis (RSA), to characterize within-sentence representational geometry. Our analysis reveals 3 key findings: (1) we observe a rank-dependence split, in which model rankings vary substantially across different similarity metrics; (2) we identify spatio-temporal alignment patterns characterized by depth-dependent alignment peaks and a pronounced increase in RSA within the 250-500 ms time window, consistent with N400-related neural dynamics; (3) we find an affective dissociation whereby negative prosody, identified using a proposed Tri-modal Neighborhood Consistency (TNC) criterion, reduces geometric similarity while enhancing covariance-based dependence. These findings provide new neurobiological insights into the representational mechanisms of Audio LLMs.

音频大模型脑电分析神经对齐语调识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。