arXiv:2508.03586cs.LGcs.AI2025-08被引 1

提出统一解释框架,让AI解释更可信且无需依赖特定模型或领域。

DeepFaith: A Domain-Free and Model-Agnostic Unified Framework for Highly Faithful Explanations

  • 基于多种可验证的忠实性度量,构建统一解释目标。
  • 在12个任务中,10项指标均优于基线方法,跨模型跨数据集表现稳定。
  • 训练后单次前向传播即可生成高保真解释,适合部署到各类AI系统。

可解释人工智能(XAI)通过模型归因方法揭示决策依据,增强对复杂系统的信任。然而,由于缺乏统一的最优解释标准,现有XAI方法缺少客观评估与优化的基准。为此,我们提出基于深度架构的忠实性解释器(DeepFaith),一个无领域、无模型依赖的统一解释框架。通过将多种广泛使用且经过验证的忠实性度量统一建模,推导出一个最优解释目标,其解能同时在多个度量下实现最优忠实性,从理论上提供解释基准。设计了一种解释学习框架,融合多种现有解释方法,通过去重和过滤构建高质量监督信号,并优化模式一致性损失与局部相关性以训练忠实解释器。训练完成后,DeepFaith仅需一次前向传播即可生成高保真解释,无需访问被解释模型。在6个模型、6个数据集上的12个多样化解释任务中,相比所有基线方法,其在10项忠实性度量上均达到最高表现,证明了方法的有效性与跨领域泛化能力。

原文摘要 · Abstract (English)

Explainable AI (XAI) builds trust in complex systems through model attribution methods that reveal the decision rationale. However, due to the absence of a unified optimal explanation, existing XAI methods lack a ground truth for objective evaluation and optimization. To address this issue, we propose Deep architecture-based Faith explainer (DeepFaith), a domain-free and model-agnostic unified explanation framework under the lens of faithfulness. By establishing a unified formulation for multiple widely used and well-validated faithfulness metrics, we derive an optimal explanation objective whose solution simultaneously achieves optimal faithfulness across these metrics, thereby providing a ground truth from a theoretical perspective. We design an explainer learning framework that leverages multiple existing explanation methods, applies deduplicating and filtering to construct high-quality supervised explanation signals, and optimizes both pattern consistency loss and local correlation to train a faithful explainer. Once trained, DeepFaith can generate highly faithful explanations through a single forward pass without accessing the model being explained. On 12 diverse explanation tasks spanning 6 models and 6 datasets, DeepFaith achieves the highest overall faithfulness across 10 metrics compared to all baseline methods, highlighting its effectiveness and cross-domain generalizability.

可解释AI忠实性统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。