研究大模型学语言时对长距离依赖的困难,发现学习速度慢而非失败。
Informational Antilocality and the Locality Bias in LLMs

- 构造无局部相关性的语言,测试模型学习能力
- 模型在不同复杂度语言上损失相似但收敛更慢
- 揭示大模型存在局部性偏差,适合研究学习机制者阅读
我们研究了基于Transformer的语言模型(LLMs)学习所谓的k-反局部语言的能力,即任意连续k个符号间无互信息的语言。通过构建k逐步增加的此类语言,发现模型在这些语言上的交叉熵损失基本相当,但随着k增大,收敛速度显著变慢。结果表明,非局部依赖更难学习,而这种偏差体现在学习速度上,而非学习成功率。这一发现为大模型中的局部性偏差提供了新证据。
原文摘要 · Abstract (English)
We consider the ability of transformer-based language models (LLMs) to learn what we call k-antilocal languages, i.e., languages that have no mutual information across any span of $k$ contiguous symbols. We construct such languages with increasing $k$, finding that LLMs trained on them achieve comparable cross-entropy loss regardless of antilocality, but converge more slowly on more antilocal languages. Our findings support the idea that non-local dependencies are more difficult to learn, but the evidence for this bias comes from learning speed rather than learning success.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。