arXiv:2606.23382cs.CLcs.AI2026-06中稿 · EMNLP

用能量模型统一预测阅读难度,效果优于传统方法。

Energy-Based Transformers as Predictors of Reading Difficulty

论文配图:Energy-Based Transformers as Predictors of Reading Difficulty
图 1 · 摘自论文原文
  • 提出基于能量的Transformer模型,连接认知记忆理论与语言处理。
  • 在三个阅读数据集上,能量值显著提升预测准确率。
  • 单层能量值即可捕捉主宾语差异,有望替代多个指标。

Transformer语言模型已成为建模人类句子处理的重要工具,其中突现度和注意力熵等指标能互补地预测阅读难度。本文首次探索能量型Transformer在计算心理语言学中的应用。该模型与关联记忆模型具有形式上的严谨联系,使语言处理研究与霍普菲尔德网络及密集关联记忆文献直接对接。在自然故事、UCL眼动追踪和自定速阅读三个阅读时间语料库中,能量度量均是阅读时间的稳健预测因子,在所有三个数据集中,其预测能力显著超越突现度和熵值。在关于关系子句加工的控制实验中,单一层的能量值成功捕捉到主语与宾语处理的不对称性。结果表明,能量度量可能同时涵盖注意力熵和突现度的影响,暗示其可作为单一统一的预测指标,取代此前所需的多个互补指标。

原文摘要 · Abstract (English)

Transformer language models have become established tools for modeling human sentence processing, with measures such as surprisal and attention entropy serving as effective predictors of reading difficulty that together capture complementary aspects of processing load. Here, we explore a related class of transformer models: energy-based transformers, which provide a principled formal link to associative memory models, bringing processing research into direct contact with the broader literature on Hopfield networks and dense associative memory. To our knowledge, this is the first exploration of an energy-based transformer measure in computational psycholinguistics. Across reading-time corpora (Natural Stories, UCL eye-tracking, UCL self-paced reading), the energy measure is a robust predictor of reading times, providing significant fit beyond surprisal and entropy in all three. In a controlled experiment on relative clause processing, energy at a single layer captures the well-known object/subject asymmetry. We find evidence that it subsumes effects attributable to both attention entropy and surprisal, suggesting that energy may serve as a single unified predictor where multiple complementary measures have previously been required.

语言模型阅读难度能量模型认知科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。