用双线性结构让模型学会反向推理,编辑更一致。
Bilinear representation mitigates reversal curse and enables consistent model editing
- 在关系图谱上训练,让模型内部表征出现双线性结构。
- 双线性结构使模型能正确推断逆向事实,避免逻辑错误。
- 适合关注模型可编辑性与逻辑一致性的研究者。
反转困境——语言模型无法从已学的“A是B”推断出未见过的“B是A”——常被视为根本缺陷。我们发现这并非模型本质缺陷,而是知识编码方式导致的产物。实验表明,从头训练合成关系知识图谱后,模型隐藏层中自然涌现出双线性关系结构,该结构缓解了反转困境,并支持对未见逆向事实的推理。关键的是,这种双线性几何结构是实现一致模型编辑的基础:单个事实的更新能正确传播至其逆向及逻辑关联关系。相比之下,缺乏此结构的模型仍受反转困境困扰,且无法泛化编辑,导致逻辑不一致。结果表明,关系知识数据集上的训练会诱导双线性内部表示的出现,从而支持模型在编辑后保持逻辑一致性。这意味着语言模型编辑的有效性不仅取决于算法选择,更取决于知识本身的底层表征几何。
原文摘要 · Abstract (English)
The reversal curse--a language model's inability to infer an unseen fact "B is A" from a learned fact "A is B"--is widely considered a fundamental limitation. We show that this is not an inherent failure but an artifact of how models encode knowledge. Our results demonstrate that training from scratch on synthetic relational knowledge graphs leads to the emergence of a bilinear relational structure within the models' hidden representations. This structure alleviates the reversal curse and facilitates inference of unseen reverse facts. Crucially, this bilinear geometry is foundational for consistent model editing: updates to a single fact propagate correctly to its reverse and logically dependent relations. In contrast, models lacking this representation suffer from the reversal curse and fail to generalize model edits, leading to logical inconsistencies. Our results establish that training on a relational knowledge dataset induces the emergence of bilinear internal representations, which in turn support language models in behaving in a logically consistent manner after editing. This suggests that the efficacy of language model editing depends not only on the choice of algorithm but on the underlying representational geometry of the knowledge itself.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。