用流匹配生成更可靠、可行动的视觉反事实解释。
LeapFactual: Reliable Visual Counterfactual Explanation Using Conditional Flow Matching
- 基于条件流匹配构建反事实生成框架,避免梯度消失问题。
- 在决策边界不一致时仍能生成分布内且准确的反事实样本。
- 适用于不可微模型和人机协作场景,提升非专家理解力。
机器学习与人工智能在医疗、科研等高风险领域日益普及,亟需既准确又可解释的模型。反事实解释通过寻找使模型预测改变的最小输入变化,提供深层洞察。然而现有方法存在梯度消失、隐空间不连续及对真实与学习决策边界对齐的过度依赖等问题。为此,我们提出LeapFactual,一种基于条件流匹配的新型反事实解释算法。该方法无需依赖可微损失函数,即使真实与学习决策边界偏离,也能生成可靠且信息丰富的反事实样本。它具备模型无关性,适用于不可微模型甚至人机协同系统(如公民科学),扩展了反事实解释的应用范围。在基准与真实数据集上的实验表明,LeapFactual生成的反事实样本准确且分布合理,可作为新训练数据提升模型性能。该方法广泛适用,有助于科学发现与非专家理解。
原文摘要 · Abstract (English)
The growing integration of machine learning (ML) and artificial intelligence (AI) models into high-stakes domains such as healthcare and scientific research calls for models that are not only accurate but also interpretable. Among the existing explainable methods, counterfactual explanations offer interpretability by identifying minimal changes to inputs that would alter a model's prediction, thus providing deeper insights. However, current counterfactual generation methods suffer from critical limitations, including gradient vanishing, discontinuous latent spaces, and an overreliance on the alignment between learned and true decision boundaries. To overcome these limitations, we propose LeapFactual, a novel counterfactual explanation algorithm based on conditional flow matching. LeapFactual generates reliable and informative counterfactuals, even when true and learned decision boundaries diverge. Following a model-agnostic approach, LeapFactual is not limited to models with differentiable loss functions. It can even handle human-in-the-loop systems, expanding the scope of counterfactual explanations to domains that require the participation of human annotators, such as citizen science. We provide extensive experiments on benchmark and real-world datasets showing that LeapFactual generates accurate and in-distribution counterfactual explanations that offer actionable insights. We observe, for instance, that our reliable counterfactual samples with labels aligning to ground truth can be beneficially used as new training data to enhance the model. The proposed method is broadly applicable and enhances both scientific knowledge discovery and non-expert interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。