arXiv:2606.19410stat.MLcs.LG2026-06

破解特征交互的混淆难题,精准分离唯一性、冗余与协同效应。

The Representational Limit of Scalar Interactions: An Interventional Decomposition

  • 通过干预式掩码推理,将特征交互分解为唯一性、冗余和协同三类机制。
  • 在表格因果模型中,交互强度恢复比基线高至411倍,显著提升结构识别能力。
  • 适用于模型可解释性分析,尤其适合关注特征协同关系的研究者。

符号化的成对交互分数本质上混淆了唯一性(U)、冗余性(R)与协同性(S)。我们在一个最简三变量异或结构因果模型上证明:忠实指标如Shapley-Taylor对每对交互返回零值,而投影型指标如Shapley Interaction将三阶效应误分配到成对标量中,造成机制混淆。本文提出Stochastic Hi-Fi,一种无需重训练的后处理可预测性分解方法,通过干预式掩码推断估计每个特征的U/R/S谱。该估计器具备精确的干预语义、有限样本蒙特卡洛界、耦合钻石采样带来的严格方差降低以及均匀有限词汇收敛性。在多个表格结构因果模型上,Stochastic Hi-Fi 恢复的结构远超标量基线(交互幅度恢复比最高达411倍)。它还能在GPT-2 IOI电路中区分冗余与协同头。在NIH ChestX-ray14数据集上,其点位游戏表现媲美GradCAM,删除AUC显著更优。

原文摘要 · Abstract (English)

Signed pairwise interaction scores fundamentally conflate uniqueness (U), redundancy (R), and synergy (S). We prove this on a minimal 3-way XOR structural causal model: faithful indices such as Shapley-Taylor return zero per pair, whereas projective indices such as Shapley Interaction spread the third-order effect into pair scalars that conflate the three mechanisms. We introduce Stochastic Hi-Fi, a post-hoc, retraining-free predictability decomposition that estimates per-feature U/R/S profiles by interventional masked inference. The estimator provides exact interventional semantics, finite-sample Monte Carlo bounds, strict variance reduction from coupled diamond sampling, and uniform finite-vocabulary convergence. Across tabular SCMs, Stochastic Hi-Fi recovers structure missed by scalar baselines (up to 411x larger interaction-magnitude recovery ratios). It also separates redundant and synergistic heads in the GPT-2 IOI circuit. On NIH ChestX-ray14, Stochastic Hi-Fi matches GradCAM on Pointing Game and improves substantially on Deletion AUC.

可解释性因果推理特征分解模型诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。