arXiv:2608.20054cs.AIcs.LG2026-08

限制模块可见信息能显著提升语言模型组合泛化能力

What You Can't See Is What You Learn: Slot-Selective Evidence Masking Favors Compositional Generalization in Shared-Genome Language-Model Societies

  • 让每个模块只看到自己负责的输入片段,而非全部输入
  • 受限可见性下模型在组合任务上平均高出20个百分点
  • 适合研究神经网络如何通过结构约束实现更优泛化

多模块神经系统通常让每个模块接触完整输入。我们测试了一种仅允许模块访问自身对应证据片段的掩码机制,是否会影响梯度训练发现的解。四单元社会共享一个冻结的预训练语言模型和一个低秩适配器,仅通过两个固定宽度的连续向量通信。在一项前瞻性封闭的自然语言函数组合任务中,我们训练了十对结构相同的受限与全局可见模型,仅注意力掩码不同。受限可见社会在两种深度下均优于全局可见对照组,9/10对中至少高出20个百分点,中位数配对优势分别为0.7648和0.6050。若切断通信,所有受限社会性能退至随机水平;事后碰撞分层分析显示,深度三模型在未训练过的仿射映射程序上仍保持0.558的准确率。六例事后选取的受限社会中,对正确回答的保留样本进行数据包干预,结果与近似值索引的中继状态一致;唯一高绩效全局模型也依赖通信,但其相同值的数据包在不同样本间不可互换。因此,受限可见性并非组合性的必要条件。在测试的种子、数据流、任务世界和训练预算下,掩码机制显著改变了训练所发现的解:事后掩码交叉实验表明两组模型均表现出掩码特异性。因受限组深度三中位准确率为0.6988(低于0.70阈值),注册实验正式失败;早期资格评估组亦无一通过。

原文摘要 · Abstract (English)

Multi-module neural systems often expose every module to the full input. We test whether a slot-selective evidence-masking regime -- restricting each module to its own evidence span -- changes which solutions gradient-based training discovers. Four-cell societies share one frozen pretrained language model and one low-rank adapter, communicating only through two model-width continuous vectors in a fixed relay. On a prospectively sealed natural-language function-composition task, we train ten matched restricted/global pairs identical except for the attention mask. Restricted-visibility societies outperform their globally visible twins by at least 20 percentage points at both depths in 9 of 10 pairs, with median paired advantages of 0.7648 and 0.6050. Cutting communication reduces every restricted society to chance, and in a post hoc collision-stratified analysis the depth-three advantage remains 0.558 on programs whose complete affine map never appeared in training. In six post hoc-selected restricted societies, packet interventions on correctly answered held-out episodes are consistent with approximately value-indexed relay states; the sole high-performing global model also requires communication, but its same-value packets are not interchangeable across episodes. Thus restricted visibility is not necessary for composition. Under the tested seeds, streams, task world, and training budget, the masking regime strongly shifted which solutions training discovered: a post hoc mask crossover finds both arms mask-native. Because the restricted mask both blocks foreign evidence and implicitly identifies each cell's assigned slot, attribution to evidence visibility alone awaits a role-marked control. The preregistered battery nevertheless formally fails because restricted-arm median depth-three accuracy is 0.6988, below the 0.70 floor; an earlier qualification cohort yielded 0/10 complete passes.

语言模型组合泛化可见性约束模块化系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。