arXiv:2505.23917cs.CVcs.AI2025-05NeurIPS

提出可解释的模型表示差异分析方法,直观揭示不同模型间的本质区别。

Representational Difference Explanations

  • 基于表示空间差异计算,直接对比两个模型的特征表达
  • 在ImageNet和iNaturalist数据集上发现有意义的表征差异与隐藏模式
  • 适合需要深度理解模型差异的研究者或模型调试人员

我们提出一种发现并可视化两个已学习表示之间差异的方法,使模型比较更加直接且可解释。我们将其称为表示差异解释(RDX),并通过比较具有已知概念差异的模型来验证该方法的有效性,结果表明它能恢复现有可解释AI(XAI)技术无法捕捉的有意义区分。将RDX应用于ImageNet和iNaturalist挑战子集上的先进模型,揭示了深刻的表征差异和细微的数据模式。尽管比较是科学分析的核心,但当前机器学习中的后处理XAI工具难以有效支持模型比较。本工作通过引入一种有效且可解释的工具,填补了这一空白。

原文摘要 · Abstract (English)

We propose a method for discovering and visualizing the differences between two learned representations, enabling more direct and interpretable model comparisons. We validate our method, which we call Representational Differences Explanations (RDX), by using it to compare models with known conceptual differences and demonstrate that it recovers meaningful distinctions where existing explainable AI (XAI) techniques fail. Applied to state-of-the-art models on challenging subsets of the ImageNet and iNaturalist datasets, RDX reveals both insightful representational differences and subtle patterns in the data. Although comparison is a cornerstone of scientific analysis, current tools in machine learning, namely post hoc XAI methods, struggle to support model comparison effectively. Our work addresses this gap by introducing an effective and explainable tool for contrasting model representations.

可解释AI表示学习模型比较

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。