将训练数据影响追踪到模型行为策略,揭示安全缺陷来源。
Symbolic Mechanistic Data Attribution: Tracing Training Influence to Learned Behavioral Policies

- 用稀疏自编码器特征建模行为政策,解析每条数据的影响路径。
- 发现宗教刻板印象等安全漏洞源于基础模型的系统性缺陷。
- 可定位无效训练对非目标特征的干扰,适合安全调优研究者。
现有数据归因方法虽能识别构建特定机制回路的训练样本,却无法解释训练数据如何塑造模型的高层行为决策。为此,我们提出符号机制数据归因(SMDA)框架,将训练样本归因于控制模型行为的可解释符号策略。SMDA通过在稀疏自编码器(SAE)特征上拟合闭式岭回归来建模目标行为,进而分析每条监督微调(SFT)样本如何通过特征激活变化(Delta_X)和输出概率变化(Delta_Y)路径影响该策略。我们对Llama-3.2-3B-Instruct的拒绝行为提取符号策略,并分析200个SFT训练样本。结果表明:(1)符号策略系数揭示了基础模型在宗教刻板印象等类别上的系统性安全缺陷;(2)按特征分解的Delta_X/Delta_Y可机制性解释有害与无害样本对某些特征产生定性差异的影响;(3)单个训练样本常引发跨特征干扰,使SMDA能识别主要作用于非预期特征的训练样本。这些结果表明,结合机制可解释性与数据归因,可实现比黑箱影响函数更精细、比人工回路分析更可扩展的诊断工具。
原文摘要 · Abstract (English)
While existing data attribution methods can identify which training examples build specific mechanistic circuits, they cannot explain how training data shapes the high-level behavioral decisions a model learns to make. To bridge this gap, we introduce Symbolic Mechanistic Data Attribution (SMDA), a framework that attributes training pairs to the interpretable symbolic policies governing model behavior. SMDA fits a closed-form Ridge regression over sparse autoencoder (SAE) features to model a target behavior, then analytically decomposes how each supervised fine-tuning example shifts that policy through feature-activation Delta_X and output-probability Delta_Y pathways. We distill a symbolic policy for refusal behavior in Llama-3.2-3B-Instruct and analyze 200 SFT training pairs. Our analysis reveals that (1) the symbolic policy's coefficients expose systematic gaps in the base model's safety behavior for categories like religious stereotyping; (2) per-feature Delta_X/Delta_Y decomposition can mechanistically explain why harmful and harmless pairs exert qualitatively different influences on certain features; and (3) individual training pairs routinely exhibit cross-feature interference, allowing SMDA to identify training pairs whose dominant effect falls on unintended features. These results demonstrate that combining mechanistic interpretability with data attribution yields a diagnostic tool that is both more fine-grained than black-box influence functions and more scalable than manual circuit analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。