arXiv:2511.16849cs.LGcs.SD2025-11

模型越能听懂声音,就越像人脑反应。

Better audio representations are more brain-like: linking model-brain alignment with performance in downstream auditory tasks

  • 用脑成像数据对比36个音频模型与大脑活动的相似性。
  • 性能越强的模型,其内部表示与脑信号越接近(相关系数>0.8)。
  • 重建音频缺失部分时,模型自然习得类脑表示,无需专门训练。

人工神经网络在模拟脑计算方面日益强大,但其下游任务表现提升是否也使内部表征更接近脑信号仍不明确。为回答此问题,我们量化了36种不同音频模型与两个独立fMRI数据集中的脑活动之间的表征对齐程度。通过体素级和成分级回归及表征相似性分析发现,近期在多种下游任务中表现优异的自监督音频模型,比以往研究的模型更能预测听觉皮层活动。为评估表征质量,我们在HEAREval基准的6个听觉任务(涵盖音乐、语音和环境声)上测试这些模型,结果揭示模型整体任务性能与脑表征对齐度之间存在强烈正相关(皮尔逊相关系数r > 0.8)。最后,我们分析了最近的音频表征模型EnCodecMAE在预训练过程中音频与脑表征相似性的演变,发现脑相似性逐步上升且早期即出现,尽管模型未显式优化此目标。这表明类脑表征可作为从自然音频数据中重建缺失信息学习过程的自然副产品。

原文摘要 · Abstract (English)

Artificial neural networks are increasingly powerful models of brain computation, yet it remains unclear whether improving their performance in downstream tasks also makes their internal representations more similar to brain signals. To address this question in the auditory domain, we quantified the alignment between the internal representations of 36 different audio models and brain activity from two independent fMRI datasets. Using voxel-wise and component-wise regression, and representation similarity analysis, we found that recent self-supervised audio models with strong performance in diverse downstream tasks are better predictors of auditory cortex activity than previously studied models. To assess the quality of the audio representations, we evaluated these models in 6 auditory tasks from the HEAREval benchmark, spanning music, speech, and environmental sounds. This revealed strong positive Pearson correlations (r > 0.8) between a model's overall task performance and its alignment with brain representations. Finally, we analyzed the evolution of the similarity between audio and brain representations during the pretraining of EnCodecMAE, a recent audio representation model. We discovered that brain similarity increases progressively and emerges early during pretraining, despite the model not being explicitly optimized for this objective. This suggests that brain-like representations can be an emergent byproduct of learning to reconstruct missing information from naturalistic audio data.

音频表征类脑模型自监督学习脑机对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。