发现大模型与大脑在抽象表征上趋同,支持深层通用特征优势。
Brain-Language Model Alignment: Insights into the Platonic Hypothesis and Intermediate-Layer Advantage
- 对比25项脑成像研究,检验模型与大脑表征一致性
- 支持柏拉图表征假说:模型越强越接近真实世界
- 揭示中间层编码更丰富通用特征,适合认知神经科学参考
近年来,神经激活与模型对齐研究兴起。本文系统回顾2023至2025年间发表的25项基于fMRI的研究,明确检验两个核心假说:(i) 柏拉图表征假说——随着模型规模扩大与性能提升,其内部表征趋于逼近真实世界的结构;(ii) 中间层优势假说——中等深度层通常编码更丰富、更具泛化能力的特征。研究结果为模型与大脑可能共享抽象表征结构提供了交叉证据,支持两项假说,并推动脑-模型对齐的进一步研究。
原文摘要 · Abstract (English)
Do brains and language models converge toward the same internal representations of the world? Recent years have seen a rise in studies of neural activations and model alignment. In this work, we review 25 fMRI-based studies published between 2023 and 2025 and explicitly confront their findings with two key hypotheses: (i) the Platonic Representation Hypothesis -- that as models scale and improve, they converge to a representation of the real world, and (ii) the Intermediate-Layer Advantage -- that intermediate (mid-depth) layers often encode richer, more generalizable features. Our findings provide converging evidence that models and brains may share abstract representational structures, supporting both hypotheses and motivating further research on brain-model alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。