arXiv:2607.10268cs.CL2026-07

探究语言局部性对模型重构能力的影响,发现模型有偏好短依赖的内在倾向。

Language Re-generation: An investigation into information locality effects on reconstruction

论文配图:Language Re-generation: An investigation into information locality effects on reconstruction
图 1 · 摘自论文原文
  • 用三种扰动方式测试GPT-2从异常语言中恢复自然语言的能力
  • 恢复出的句法依赖长度更短,体现模型对局部结构的固有偏好
  • 长句在结构破坏下恢复失败,适合研究模型归纳偏置的读者必看

信息局部性指语法相关词语倾向于邻近出现,影响人类语言处理与语言模型学习。现有研究关注模型能否学会不可能语言,但尚不清楚其能否从中恢复自然语言,以及这揭示了何种归纳偏置。本文通过重建框架补充可学习性研究:微调预训练于不可能语言的GPT-2模型,从三类扰动输入中恢复自然英语。结果表明,恢复出的结构具有更短的依赖长度,与无约束生成中的局部偏好一致,提供了架构偏置的定量证据,而可学习性实验无法揭示此现象。恢复难度随局部性破坏程度增加。结构恢复(依赖三元组F1)与表面恢复(精确匹配)解耦,流畅性与忠实重构在全局打乱下亦解耦。句长进一步调节表现:当局部结构保留时长句利于恢复,但在全局打乱下导致完全崩溃。最后,恢复难度与可学习性难度在不同扰动类型间高度一致,表明信息局部性是二者共享的约束条件。

原文摘要 · Abstract (English)

Information locality, the tendency for syntactically related words to appear close together, shapes both human language processing and language model learning. While prior work has examined whether language models can acquire impossible languages, it remains unclear whether they can recover natural language from such input and what this reveals about their inductive biases. We address this by complementing learnability-based approaches with a reconstruction framework: fine-tuning GPT-2 models pre-trained on impossible languages to reconstruct natural English from three perturbation types. Our findings show that the recovered structures exhibit shorter dependency lengths than the original text, mirroring the locality preference observed in unconstrained language model generation and providing a quantitative signature of an architectural bias that learnability experiments alone do not reveal. Recovery difficulty increases with the degree of locality disruption. Structural recovery (dependency Triple F1) dissociates from surface recovery (Exact Match), while fluency dissociates from faithful reconstruction under global shuffling. Sentence length further modulates performance: longer sentences facilitate recovery when local structure is preserved but lead to complete collapse under global shuffling. Finally, recovery difficulty tracks learnability difficulty across perturbation types, suggesting that information locality is the shared constraint governing both.

语言模型归纳偏置句法恢复局部性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。