Mamba模型的处理时长与人类阅读时间高度吻合,揭示了语言模型与认知机制的潜在联系。
Timesteps of Mamba Align with Human Reading Times

- Mamba每词处理时间由输入动态决定,模拟人类实时阅读过程。
- 该时长可显著预测人类阅读时间,超越传统语言模型的预测能力。
- 适合研究认知科学、语言处理与神经网络动态交互的学者。
本研究揭示了主流状态空间语言模型Mamba中每个词的处理时间与人类阅读时间之间的对齐关系。在Mamba中,每一层的递归状态转移具有概念上的持续时间,即随输入动态调整的离散时间步长Δ_t。基于自然阅读数据集,我们发现Mamba的每词时间步长是人类阅读时间的重要预测因子,即使在控制GPT-2意外性等已知变量后仍显著。通过形式化分析Mamba的架构与内部动态,我们进一步提出,Mamba可作为观察人类实时语言处理的新视角,尤其能揭示各模块如何权衡短期与长期信息保留,以及噪声如何影响连续记忆表示。代码已公开。
原文摘要 · Abstract (English)
This study demonstrates an alignment of per-word processing time in a popular state-space language model Mamba and human readers. In Mamba, the recurrent state transition at each layer conceptually takes some duration of time, the discretization timestep $Δ_t$, determined dynamically in response to the input. Using a naturalistic reading dataset, we show that the per-word timestep from Mamba is a significant predictor of human reading times, and remains significant even when known predictors such as GPT-2 surprisal are controlled for. We further suggest, through formal analysis of Mamba's architecture and internal dynamics, that Mamba can serve as a new, valuable lens to look at human real-time language processing with ever-updated memory, because it allows us to look at how each module (layer) weighs short- and long-term information retention, and how noise may interact with dynamic, continuous memory representation. Code is available online.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。