分解模型扰动响应中的证据、矛盾与脆弱性,更精准解释决策依据。
Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses
- 将响应分解为证据、矛盾和脆弱性三部分,保持总量不变
- 96.4%的案例中DECAF组件与实际行为一致,远超仅看响应大小
- 短轨迹可大幅降低计算资源,适合大规模模型解释
扰动方法通过测量输入变化时预测的改变来解释模型决策,但响应幅度仅反映模型反应强度,无法说明其意义。相同幅度可能支持、反对或在路径中强而终点消失。我们追踪成对输入逐步揭示过程中对比的变化,用最终对比解释轨迹。提出DECAF(证据、矛盾与脆弱性分解),将一致、对立及终点归零的响应分别分配给证据E、矛盾C、脆弱性F。该分解保持原始幅度总和,即Abs = E + C + F,且在终点相对公理下唯一。在控制视觉与表格设置中,三成分独立对应行为。72模型ImageNet-9审计中,响应幅度相近但行为不同的案例,最大DECAF成分与观察行为匹配率达96.4%,远超幅度单独判断的35.0%。仅改变揭示路径使总响应提升近80%,证据基本不变,脆弱性增长超4倍。在FunnyBirds与ImageNet-1k上,短前向DECAF轨迹优于现有通用归因基线。在10亿参数DINOv2模型上,短轨迹以4.75倍更低的墙时间与2.36倍更低峰值内存,达到强梯度基线效果。
原文摘要 · Abstract (English)
Perturbation methods explain model decisions by measuring prediction changes under altered inputs, but response magnitude tells us only how much a model reacts, not what that reaction means. The same magnitude can support the final factual-counterfactual difference, oppose it, or arise strongly along the perturbation path yet vanish at the endpoint. We therefore track how the contrast develops as paired inputs are progressively revealed, using the final contrast to interpret the trajectory. We introduce DECAF (Decomposition of Evidence, Contradiction, And Fragility), which routes aligned, opposed, and endpoint-null responses into evidence E, contradiction C, and fragility F. The decomposition preserves ordinary magnitude exactly, Abs = E + C + F, and is unique under endpoint-relative axioms. Across controlled vision and tabular settings, the three components track independently measured behavior. In a 72-model ImageNet-9 audit, we compare cases with nearly identical response magnitude but different independently measured behaviors. The largest DECAF component agrees with an observed behavior in 96.4% of cases, compared with 35.0% for magnitude alone. Changing only the reveal path increases total response by nearly 80%, yet evidence barely changes while fragility grows by more than 4x. On FunnyBirds and ImageNet-1k, short forward-only DECAF trajectories outperform the tested general-purpose attribution baselines. On a 1B-scale DINOv2 model, a short trajectory matches a strong gradient-based baseline with 4.75x lower wall time and 2.36x lower peak memory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。