通过模拟大脑拓扑结构,让语音模型生成更接近真实脑区的表征。
Topographic Constraints Shape Brain-Like Component Structure in Auditory Models
- 在神经网络中加入空间拓扑约束,使相邻单元响应相似。
- 模型内部表征更紧凑,且与人脑ECoG数据中的成分高度匹配。
- 适合研究神经表征机制或想提升模型生物合理性的研究人员。
若拓扑是大脑的基本特征,则它不仅影响神经元的空间分布(即解释脑图谱),还应影响神经群体内信息的组织方式。人类听觉皮层为检验后一种观点提供了有力但此前未充分利用的测试场景。通过fMRI和ECoG测得的神经反应可分解为对应于语音、音乐、歌曲等声音类别的可解释成分,揭示了声音信息在大脑中的划分方式。本文探讨在音频神经网络训练中引入拓扑约束是否能使其内部表征更贴近脑中观察到的成分结构。为此,我们提出一类新型拓扑听觉模型——TopoAudio,其包含布线长度约束,并促使二维皮层网格上相邻单元发展出相似的响应调谐特性。尽管增加这些约束,TopoAudio在标准语音与环境声分类任务上表现与非拓扑模型相当,且在预测人脑fMRI响应方面也表现一致。关键的是,拓扑模型产生更紧凑的内部表征,其推断出的成分与人脑ECoG记录的成分更加吻合。结果初步表明,拓扑结构是生成生物合理内部表征的通用机制。更广泛而言,成分级对齐为检验拓扑是否重塑模型种群响应以更好匹配神经记录中的表征结构提供了补充方法。
原文摘要 · Abstract (English)
If topography is a fundamental feature of the brain, it should influence both how neurons are arranged in space (i.e. explain brain maps) and how information is structured within the neural population. The human auditory cortex provides a strong, but previously underused test for the latter idea. Neural responses measured with both fMRI and ECoG can be decomposed into interpretable components corresponding to sound categories such as speech, music, and song, offering a view of how sound information is partitioned in the brain. Here we ask whether introducing topographic constraints into the training of audio neural network models shapes their internal representations to better match the component structure observed in the brain. To address this question, we introduce a new class of topographic auditory models, TopoAudio, which incorporate wiring-length constraints and encourage nearby units on a two-dimensional cortical sheet to develop similar response tuning. Despite these additional constraints we find that TopoAudio achieves comparable performance on standard speech and environmental sound classification tasks to standard non-topographic models and matches them in predicting human fMRI responses. Crucially however, topographic models develop more compact internal representations, and their inferred components align more closely with those derived from human ECoG recordings. These results provide initial evidence that topography offers a general mechanism for producing biologically aligned internal representations in artificial neural networks. More broadly, component-level alignment provides a complementary way for testing whether topography reshapes population responses in models to better match the representational structure observed in neural recordings. Our project page: https://topoaudio.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。