arXiv:2511.06944cs.CVcs.AI2025-11AAAI

让模型预测与解释一起优化,提升可解释性与泛化能力。

From Attribution to Action: Jointly ALIGNing Predictions and Explanations

  • 迭代训练分类器与掩码生成器,联合优化预测与解释。
  • 在VLCS和Terra Incognita上超越6个基线,跨域性能更优。
  • 生成的解释更完整且充分,适合需要可信AI的场景。

解释引导学习(EGL)在计算机视觉任务中展现出将模型预测与可解释推理对齐的潜力。然而,现有方法多依赖外部标注或启发式分割来监督解释,这些信号往往噪声大、不精确且难以扩展。本文提供了实证与理论证据,表明低质量监督信号反而会损害模型性能。为此,我们提出ALIGN框架,通过迭代方式联合训练分类器与掩码生成器。掩码生成器学习生成软性的、任务相关的显著区域掩码,而分类器则同时优化预测准确率及显著图与掩码之间的对齐度。利用高质量掩码作为指导,ALIGN在可解释性与泛化能力上均表现更优。在两个领域泛化基准(VLCS与Terra Incognita)上的实验表明,ALIGN在分布内与分布外设置下均持续优于六个强基线。此外,其生成的解释在充分性与完整性方面也显著更优,验证了该方法在生成准确且可解释模型方面的有效性。

原文摘要 · Abstract (English)

Explanation-guided learning (EGL) has shown promise in aligning model predictions with interpretable reasoning, particularly in computer vision tasks. However, most approaches rely on external annotations or heuristic-based segmentation to supervise model explanations, which can be noisy, imprecise and difficult to scale. In this work, we provide both empirical and theoretical evidence that low-quality supervision signals can degrade model performance rather than improve it. In response, we propose ALIGN, a novel framework that jointly trains a classifier and a masker in an iterative manner. The masker learns to produce soft, task-relevant masks that highlight informative regions, while the classifier is optimized for both prediction accuracy and alignment between its saliency maps and the learned masks. By leveraging high-quality masks as guidance, ALIGN improves both interpretability and generalizability, showing its superiority across various settings. Experiments on the two domain generalization benchmarks, VLCS and Terra Incognita, show that ALIGN consistently outperforms six strong baselines in both in-distribution and out-of-distribution settings. Besides, ALIGN also yields superior explanation quality concerning sufficiency and comprehensiveness, highlighting its effectiveness in producing accurate and interpretable models.

可解释AI联合优化领域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。