arXiv:2507.16080q-bio.NCcs.SD2025-07被引 3

揭示语音模型如何编码大脑对声音的响应,助力构建更可解释的脑机模型。

Interpretable Embeddings of Speech Enhance and Explain Brain Encoding Performance of Audio Models

  • 用六类可解释特征建模语音模型与大脑反应的关联。
  • 大模型能捕捉低层特征外的脑相关语义,且能力随规模提升。
  • 提示通过融合可解释特征,提升模型可解释性与精度。

语音基础模型(SFMs)被视为人类语音感知的强大计算模型,但其表征本质为黑箱,难以理解其与大脑反应对齐的机制。为此,我们基于六类可解释特征(梅尔频谱图、伽伯滤波器组、语音存在性、音素、句法、语义)及三类先进SFMs(Whisper、HuBERT、WavLM)的上下文嵌入,构建线性编码模型,量化这些特征类别在电皮质图(ECoG)响应方差中的共享程度。方差分解分析显示:第一,SFMs与大脑的对齐主要源于其学习和编码简单可解释语音特征的能力;第二,不同层间存在低级与高级特征编码的系统性权衡;第三,大模型能学习到无法由低级特征解释的脑相关语义,且该能力随模型规模和上下文长度增加而增强。研究结果表明,通过在SFMs嵌入中引入可解释特征,可构建更可解释、准确且高效的脑编码模型。

原文摘要 · Abstract (English)

Speech foundation models (SFMs) are increasingly hailed as powerful computational models of human speech perception. However, since their representations are inherently black-box, it remains unclear what drives their alignment with brain responses. To remedy this, we built linear encoding models from six interpretable feature families: mel-spectrogram, Gabor filter bank features, speech presence, phonetic, syntactic, and semantic features, and contextualized embeddings from three state-of-the-art SFMs (Whisper, HuBERT, WavLM), quantifying electrocorticography (ECoG) response variance shared between feature classes. Variance-partitioning analyses revealed several key insights: First, the SFMs' alignment with the brain can be mostly explained by their ability to learn and encode simple interpretable speech features. Second, SFMs exhibit a systematic trade-off between encoding of brain-relevant low-level and high-level features across layers. Finally, our results show that SFMs learn brain-relevant semantics which cannot be explained by lower-level speech features, with this capacity increasing with model size and context length. Together, our findings suggest a principled approach to build more interpretable, accurate, and efficient encoding models of the brain by augmenting SFM embeddings with interpretable features.

语音模型脑机接口可解释性深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。