arXiv:2601.13548cs.LG2026-01被引 6

通过逆向推导数据,让模型按需生成特定内部结构。

Patterning: The Dual of Interpretability

  • 用敏感度分析反推训练数据调整方式,引导模型形成目标结构。
  • 在小语言模型中实现结构加速或延迟形成,如归纳电路。
  • 可选择模型学习的算法,即使所有方案训练准确率相同。

机制可解释性旨在通过逆向解析神经网络内部结构来理解其泛化能力。本文提出‘模式构建’作为其对偶问题:给定期望的泛化形式,反推应如何设计训练数据以实现该结果。方法基于敏感度(susceptibilities),衡量可观测量后验期望值对数据分布微小扰动的响应。通过反演这一线性响应关系,可得到使模型趋向目标内部状态的数据干预策略。我们在一个小语言模型中验证了该方法,发现沿主敏感度方向重加权训练数据,能加速或延迟结构(如归纳电路)的形成。在合成括号匹配任务中,多个算法均达完美训练准确率,但通过靶向各解法的局部学习系数,可实现对模型学习算法的选择。结果表明,用于解析内部结构的同一数学框架也可反向用于主动构造结构。

原文摘要 · Abstract (English)

Mechanistic interpretability aims to understand how neural networks generalize beyond their training data by reverse-engineering their internal structures. We introduce patterning as the dual problem: given a desired form of generalization, determine what training data produces it. Our approach is based on susceptibilities, which measure how posterior expectation values of observables respond to infinitesimal shifts in the data distribution. Inverting this linear response relationship yields the data intervention that steers the model toward a target internal configuration. We demonstrate patterning in a small language model, showing that re-weighting training data along principal susceptibility directions can accelerate or delay the formation of structure, such as the induction circuit. In a synthetic parentheses balancing task where multiple algorithms achieve perfect training accuracy, we show that patterning can select which algorithm the model learns by targeting the local learning coefficient of each solution. These results establish that the same mathematical framework used to read internal structure can be inverted to write it.

可解释性模型控制数据干预神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。