arXiv:2607.01033cs.LG2026-07被引 2

不同训练方法让模型行为可解释性差异巨大,现有测试用模型可能不靠谱。

The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology

论文配图:The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology
图 1 · 摘自论文原文
  • 用7种训练法构建54个测试模型,比较解释性差异。
  • 真实整合训练的模型反而更难解释,传统方法太简单。
  • 提醒研究者:别再拿现成测试模型当标准了。

模型生物(MOs)是通过后处理监督微调等方法训练出的、表现出异常行为的语言模型,常被用来评估白盒可解释性技术。已有研究表明这些模型容易被识别出隐藏行为,但近期研究指出这类训练方式可能使解释任务过于简单。本文构建了基于 54 个 OLMo2-1B 与 gemma-3-1b-it 模型的多种 MO 变体,采用包括标准后处理 SFT、DPO 以及更真实的将目标数据融入 OLMo 后训练阶段 DPO 的七种训练方法。使用激活预言机、激活引导、逻辑透镜和稀疏自编码器进行评测。结果表明:(i)MO 解释性强烈依赖于训练目标、目标行为、模型架构及数据生成流程;(ii)即使控制行为强度差异,仍存在显著变异性;(iii)更贴近真实场景的“集成训练”方法通常得到的 MO 更难解释。结论质疑当前 MO 作为可解释性代理的有效性。

原文摘要 · Abstract (English)

Model organisms (MOs) - language models trained to exhibit undesired or unnatural behaviours - are frequently used as testbeds for evaluating white-box interpretability techniques. Current MOs are typically constructed via post-hoc supervised fine-tuning (SFT) on behavioural transcripts or synthetic documents. Prior research has shown that interpretability methods can easily identify hidden behaviours in these MOs. However, recent work suggests that such post-hoc training methods may make interpretability unrealistically easy. We investigate this claim by constructing a suite of 54 $\verb|OLMo2-1B|$- and $\verb|gemma-3-1b-it|$-based MOs trained with seven different techniques, including standard post-hoc SFT, post-hoc DPO, and more realistic integration of MO data into the OLMo post-training DPO phase. We use these MO variants to benchmark activation oracles, activation steering, logit lens, and sparse autoencoders. Our findings show that (i) MO interpretability depends strongly on training objective, target behaviour, model architecture, and training data generation pipeline; (ii) substantial variance remains even after controlling for differences in the strength of target behaviour expression; and (iii) our more realistic $\textit{integrated training}$ often yields less interpretable MOs than standard post-hoc methods. Our results cast substantial doubt on the validity of current MOs as interpretability proxies.

可解释性模型生物训练方法验证可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。