揭示大模型在可逆事实学习上的失败源于概念绑定难题。
Is the Reversal Curse a Binding Problem? Uncovering Limitations of Transformers from a Basic Generalization Failure
- 提出概念表示不一致与纠缠是导致错误的关键原因
- 首次在不依赖特殊数据增强下打破反转诅咒
- 引入记忆层提升概念解耦,增强模型泛化能力
尽管大语言模型表现卓越,但在可逆事实关联学习上存在基本泛化失败,即反转诅咒。本文推测该现象源于认知科学、神经科学和人工智能中的长期绑定问题。具体假设变压器模型在概念绑定上的局限性主要来自概念表示的不一致性和纠缠性。通过一系列实验验证了这一猜想。研究进一步提出基于JEPA的模型设计,首次在不使用特殊数据增强或非因果掩码的情况下突破反转诅咒;并通过引入支持解耦概念表示的特殊记忆层,进一步提升泛化性能。本工作揭示了构建具备系统性概念绑定能力模型的根本挑战,推动减少对人工引导的依赖。
原文摘要 · Abstract (English)
Despite their impressive capabilities, LLMs exhibit a basic generalization failure known as the Reversal Curse, where they struggle to learn reversible factual associations. Understanding why this occurs could help identify weaknesses in current models and advance their generalization and robustness. In this paper, we conjecture that the Reversal Curse in LLMs is a manifestation of the long-standing binding problem in cognitive science, neuroscience and AI. Specifically, we hypothesize two primary causes of the Reversal Curse stemming from transformers' limitations in conceptual binding: the inconsistency and entanglements of concept representations. We perform a series of experiments that support these conjectures. Our exploration leads to a model design based on JEPA (Joint-Embedding Predictive Architecture) that for the first time breaks the Reversal Curse without side-stepping it with specialized data augmentation or non-causal masking, and moreover, generalization could be further improved by incorporating special memory layers that support disentangled concept representations. Our research opens up the broader fundamental challenge of designing models capable of learning systematic conceptual binding with less human scaffolding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。