arXiv:2506.23074cs.CVcs.CR2025-06ICCV被引 5

通过反事实解耦注意力,提升模型溯源在开放世界中的泛化能力

Learning Counterfactually Decoupled Attention for Open-World Model Attribution

  • 基于反事实推理解耦注意力与源模型关联,消除干扰偏差
  • 在未见过的新攻击上显著优于现有方法,性能提升明显
  • 适合关注模型可解释性与开放世界溯源的研究者

本文提出一种反事实解耦注意力学习(CDAL)方法,用于开放世界模型溯源。现有方法依赖手工设计的区域划分或特征空间,易受虚假统计相关性干扰,在开放世界新攻击场景下表现不佳。CDAL 显式建模注意力痕迹与源模型归属之间的因果关系,反事实地将判别性模型特异性特征与混淆性源偏差解耦,以实现更准确的比较。由此得到的因果效应可量化注意力图质量,促使网络最大化该效应,从而捕捉泛化到未知源模型的本质生成模式。在多个公开的开放世界模型溯源基准上,本方法以极小计算开销持续显著提升当前最优模型性能,尤其在未见过的新攻击上优势突出。

原文摘要 · Abstract (English)

In this paper, we propose a Counterfactually Decoupled Attention Learning (CDAL) method for open-world model attribution. Existing methods rely on handcrafted design of region partitioning or feature space, which could be confounded by the spurious statistical correlations and struggle with novel attacks in open-world scenarios. To address this, CDAL explicitly models the causal relationships between the attentional visual traces and source model attribution, and counterfactually decouples the discriminative model-specific artifacts from confounding source biases for comparison. In this way, the resulting causal effect provides a quantification on the quality of learned attention maps, thus encouraging the network to capture essential generation patterns that generalize to unseen source models by maximizing the effect. Extensive experiments on existing open-world model attribution benchmarks show that with minimal computational overhead, our method consistently improves state-of-the-art models by large margins, particularly for unseen novel attacks. Source code: https://github.com/yzheng97/CDAL.

模型溯源因果推理注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。