语言模型看似能预测大脑反应,但实则多为噪音干扰,不能算真正对齐。
Do Language Models Align with Brains? Prediction Scores Are Not Enough
- 用严格对照框架检验模型与大脑的匹配度
- 414组预测、2304组关系等数据均被控因解释
- 强调不能仅凭预测分数就断言模型对齐大脑
脑-语言模型对比常将神经预测分数视为模型捕捉大脑语言计算的证据。本文使用L-PACT框架,评估预测性、关系性、机制剥离及可靠性边界证据,覆盖414组预测控制行、2304组关系特征行、4320组机制剥离行、420组脑-脑上限行和146组综合决策行。实验显示,所有真实模型行均未通过预测、关系、机制剥离或操作图灵限制的可靠性关卡;全部146个综合判断结果均被控制变量解释。尽管在敏感性测试中,脑-脑一致性等信号可产生假阳性,但实际模型表现无法排除基线干扰。结论:当前语言模型表示未满足L-PACT对齐标准,所谓正向结果应归为可控解释而非结构对齐。
原文摘要 · Abstract (English)
Brain-language model comparisons often interpret neural prediction scores as evidence that model representations capture brain-relevant language computation. We asked whether language models align with brains, and whether prediction scores are enough to support that claim, using L-PACT, a source-audited framework that evaluates predictive, relational, mechanism-stripping, and reliability-bounded evidence. Across primary naturalistic language neural datasets and derived language-model representations, L-PACT compared real model features with nuisance baselines and severe controls, tested whether model-to-brain profiles reproduced brain-to-brain patterns, recomputed held-out scores after mechanism stripping, and normalized evidence against brain-brain ceilings. The locked analysis set contains 414 predictive-control rows, 2304 relational profile rows, 4320 mechanism-stripping rows, 420 brain-brain ceiling rows, and 146 integrated decision rows. Assay-sensitivity checks showed that brain-brain reliability, brain-as-model run-to-run relational profiles, independent low-level neural and WAV-derived acoustic-envelope gates, and a deterministic implanted-signal simulation can produce positive evidence when expected. Nevertheless, no real model row passed the predictive, relational, mechanism-stripping, or operational Turing-bounded reliability gates; all 146 integrated rows were control-explained. Less stringent single-criterion rules would have counted raw positive predictive, relational, stripping-delta, and ceiling-normalized effects, but L-PACT downgraded them because controls explained the apparent evidence. In the analyzed derived artifact set, the tested language-model representations do not satisfy L-PACT alignment gates; apparent positives are converted into an auditable control-explained taxonomy rather than treated as structural alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。