arXiv:2501.08537cs.CLcs.LG2025-01TPAMI被引 19

控制复杂度能让Transformer学会推理而非死记硬背。

Complexity Control Facilitates Reasoning-Based Compositional Generalization in Transformers

  • 通过遮蔽信息通路,发现复杂度控制影响模型学习方式。
  • 低复杂度偏差的模型更易获得泛化性推理能力。
  • 适用于希望提升模型泛化能力的研究者。

Transformer在多种任务中表现优异,但在组合性问题上的表现仍存争议。本研究探究了Transformer在组合任务中的内部机制。结果表明,复杂度控制策略显著影响模型是否学习到可外推的原始规则(推理型解法)或仅依赖记忆映射(记忆型解法)。通过遮蔽模型信息通路并使用多种复杂度度量,揭示了不同解法对应的不同内部工作机制。进一步分析发现,推理型解法具有更低的复杂度偏差,这与已知的神经元凝聚现象一致。该低复杂度偏差被认为是实现推理规则学习的关键因素。研究在图像生成和自然语言处理等多个真实数据集上验证了结论,确认了其广泛适用性。

原文摘要 · Abstract (English)

Transformers have demonstrated impressive capabilities across various tasks, yet their performance on compositional problems remains a subject of debate. In this study, we investigate the internal mechanisms underlying Transformers' behavior in compositional tasks. We find that complexity control strategies significantly influence whether the model learns primitive-level rules that generalize out-of-distribution (reasoning-based solutions) or relies solely on memorized mappings (memory-based solutions). By applying masking strategies to the model's information circuits and employing multiple complexity metrics, we reveal distinct internal working mechanisms associated with different solution types. Further analysis reveals that reasoning-based solutions exhibit a lower complexity bias, which aligns with the well-studied neuron condensation phenomenon. This lower complexity bias is hypothesized to be the key factor enabling these solutions to learn reasoning rules. We validate these conclusions across multiple real-world datasets, including image generation and natural language processing tasks, confirming the broad applicability of our findings.

Transformer推理泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。