用脑电图分析语言模型的预测行为是否像人一样真实。
Encoding EEG Signals to Examine Human-Like Next-Word Prediction Behaviour in Language Models

- 用预测准确率和意外度构建人类与模型的脑电预测回归器。
- 只有意外度能有效预测脑电反应,尤其对语义丰富的词。
- 大模型不等于更像人,规模扩大未必提升认知一致性。
语言模型(LMs)在给定上下文时擅长预测下一个词,人类阅读也具有类似可预测性。神经科学表明,词的可预测性会影响脑电图(EEG)记录的毫秒级脑响应。尽管先进语言模型在下一个词预测任务上的准确率已接近人类水平,但问题在于:高预测准确率是否意味着模型真正捕捉了人类阅读理解的认知信号?为此,我们基于两种信息量度——顶1预测和意外度——构建人类与模型的回归器,用于预测从EEG记录中提取的事件相关电位(ERP),反映阅读过程中的不同认知阶段。结果表明,只有意外度与语言处理相关的ERP存在显著关联,尤其在语义内容丰富的开放类词上。研究还挑战了“模型规模越大,越接近人类语言处理”的假设,提示单纯扩大参数和计算资源并不保证认知拟合度提升。
原文摘要 · Abstract (English)
Language models (LMs) are trained to excel at predicting the next word in the sequence given prior context, and humans also share this predictability in reading comprehension. Neuroscience research reveals that next-word predictability influences brain response, as recorded at millisecond resolution using electroencephalography (EEG). While our evidence indicates that advanced LMs achieve accuracies closely aligned with human performance at the next-word prediction task, this raises the question: Does higher prediction accuracy necessarily mean that these models adequately capture the cognitive signals associated with human reading comprehension? Here, we generate regressors for both humans and LMs based on two information measures, including top-1 prediction and surprisal, to predict event-related potential (ERP) elicited from EEG recordings which reflect different stages of cognitive processing during reading. We argue that modelling ERP patterns offers fine-grained analysis of the cognitive plausibility of various LMs during reading. Our results indicate that only surprisal potentially correlates with language-processing ERPs, especially for open-class words with high semantic content. Moreover, our findings challenge the assumption that scaling LMs with increased parameters and computational budgets will consistently lead to improved convergence with human-like linguistic processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。