用内部机制预测模型在未见数据上的行为,验证了可解释性新方向。
Can Interpretation Predict Behavior on Unseen Data?
- 基于分布内数据的注意力模式预测模型在分布外的表现。
- 成功预测了上百个Transformer在不同泛化规则下的行为差异。
- 即使因果关系不明确,观测分析仍能有效预测行为,适合可信AI研究者。
可解释性研究通常预测模型对特定机制干预的响应,但能否预测其在未见输入数据上的表现?我们提出并验证了这一新目标:利用模型内部状态预测其在分布外(OOD)数据上的行为。我们在数百个Transformer上训练了简单合成任务,这些任务中分布内准确率完美,但存在多种可能的分布外泛化规则。我们发现,仅从分布内数据中观察到的注意力模式,即可成功预测每个模型在分布外数据上遵循的具体规则。实验分离了解释的机制忠实性与预测价值:消融研究表明,某些内部模式会抑制而非支持其所预测的规则,说明即使缺乏简单的因果关联,观测分析也能准确预测行为。该成果为可解释性提出了新范式:通过理解模型内部机制来预测行为并评估分布偏移下的可靠性。
原文摘要 · Abstract (English)
Interpretability research often predicts model responses to targeted mechanistic interventions. But can we predict responses to unseen input data? We propose and demonstrate this alternate objective by using model internals to predict their out-of-distribution (OOD) behavior. We train hundreds of Transformers on simple synthetic tasks, where perfect in-distribution accuracy is compatible with multiple OOD generalization rules. We successfully use attention patterns -- observed only on in-distribution data -- to predict which rule each model follows on OOD data. Our experiments decouple the mechanistic faithfulness of our interpretation from its predictive value; ablations reveal such internal patterns can suppress rather than support the rule they predict, showing observational analysis can forecast behavior even when causal analysis fails to support a simple cause-effect link. Our findings are a proof-of-concept for a new interpretability objective: understanding model internals to predict behavior and assess reliability under distribution shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。