arXiv:2505.11770cs.LGcs.AI2025-05ICML被引 12

用内部因果机制预测模型在分布外数据的表现

Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors

  • 通过因果变量模拟和值探测预测输出正确性
  • 分布外场景下AUC-ROC表现优于传统方法
  • 适合关注模型可解释性与鲁棒性的研究者

可解释性研究提供了多种识别神经网络内部抽象机制的方法。这些方法能否用于预测模型在分布外样本上的行为?本文给出肯定回答。在符号操作、知识检索和指令遵循等多种语言建模任务中,我们发现对正确性预测最稳健的特征是那些在模型行为中扮演独特因果角色的特征。为此,我们提出两种基于因果机制的预测方法:反事实模拟(检查关键因果变量是否实现)和值探测(利用这些变量的取值进行预测)。两者在分布内均取得高AUC-ROC,在分布外设置下显著优于依赖因果无关特征的方法,而分布外预测恰是实际应用中最关键的需求。本工作揭示了语言模型内部因果分析的新且重要的应用场景。

原文摘要 · Abstract (English)

Interpretability research now offers a variety of techniques for identifying abstract internal mechanisms in neural networks. Can such techniques be used to predict how models will behave on out-of-distribution examples? In this work, we provide a positive answer to this question. Through a diverse set of language modeling tasks--including symbol manipulation, knowledge retrieval, and instruction following--we show that the most robust features for correctness prediction are those that play a distinctive causal role in the model's behavior. Specifically, we propose two methods that leverage causal mechanisms to predict the correctness of model outputs: counterfactual simulation (checking whether key causal variables are realized) and value probing (using the values of those variables to make predictions). Both achieve high AUC-ROC in distribution and outperform methods that rely on causal-agnostic features in out-of-distribution settings, where predicting model behaviors is more crucial. Our work thus highlights a novel and significant application for internal causal analysis of language models.

可解释性因果推理分布外

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。