arXiv:2608.23851cs.CL2026-08

用记忆检索提升模型对生僻词的语法判断能力

Does Episodic Memory Help Close the Lexical Frequency Gap in Sensitivity to Syntactic Contrasts? A Test Using Retrieval-Augmented Language Models

  • 引入检索增强机制模拟海马体记忆,存储具体语言实例
  • 生僻词语法判断准确率提升,频率差距显著缩小
  • 适合研究语言习得、神经认知模型与低频语言现象

语法知识及其测试通常被认为不受词汇频率影响,但基于神经网络的语言模型对此高度敏感。本文基于互补学习系统理论,假设海马体式的情景记忆可帮助克服频率偏差,通过快速编码与检索特定语言经验来弥补参数化表示的不足。我们采用检索增强语言模型(即基于k近邻的模型,将显式实例存储与参数模型结合)进行验证,测试其是否能缓解普通语言模型在语法对比测试中对高频与低频词汇的性能差距。使用频率分层的语法对比测试集,结果表明:检索增强可缩小高频与低频词之间的表现差距,且该效果在不同句法现象及预训练于儿童语料和大规模数据的模型中均成立。此外,结构信息对有效检索至关重要,仅靠语义相似性则收效甚微。尽管这些结果初步支持假设,但频率差距仍未完全消除。未来方向包括:优先重加权检索实例、改进结构表示与检索策略、灵活配置存储与检索机制。

原文摘要 · Abstract (English)

Grammatical knowledge and how it is empirically tested are typically considered robust to the frequency of the lexical items in the expressions. However, neural network-based models of grammaticality exhibit high sensitivity to lexical frequency. We draw upon Complementary Learning Systems theory to test the hypothesis that robustness to lexical frequency can arise via a hippocampal episodic memory mechanism, which enables rapid encoding and retrieval of specific experiences and allows learners to leverage them when processing rare patterns. We use retrieval-augmented language models as an instantiation of such an episodic memory mechanism (specifically, $k$-nearest-neighbor language models that augment parametric models with explicit instance storage), and test whether this augmentation helps close the lexical frequency gap that vanilla language models exhibit in syntactic contrast tests. Using syntactic contrasts with frequency-stratified test items, we find that retrieval augmentation narrows the performance gap between high- and low-frequency items, consistent with episodic memory compensating for weak parametric representations. This benefit is consistent across different syntactic phenomena and across models pretrained on child-realistic and large-scale data. Additionally, we show that structural information is critical for effective retrieval, whereas semantic similarity alone provides little benefit. While these are promising proof-of-concept results supporting our hypothesis, the frequency gap is narrowed rather than fully closed. Based on our analyses, we propose preferential reweighting of retrieved instances, better representations and retrieval strategies for structural information, and flexible configurations of storage and retrieval as promising future directions for improving the implementation of episodic memory in language models.

语言模型记忆机制语法敏感性检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。