arXiv:2505.06123cs.LGcs.AI2025-05TPAMI被引 2

用可解释AI拆解Wasserstein距离,看清数据差异的根源

Wasserstein Distances Made Explainable: Insights Into Dataset Shifts and Transport Phenomena

  • 通过可解释AI将距离归因到具体数据组或特征
  • 在多种数据集和距离设定下保持高精度
  • 适合关注数据分布变化或偏差分析的研究者

Wasserstein距离为比较数据分布提供了强大框架,可用于分析随时间变化的过程或检测数据内部的不均匀性。然而,仅计算Wasserstein距离或分析对应的运输计划(耦合)可能不足以理解导致高或低距离的具体因素。本文提出一种基于可解释AI的新方法,能够高效且准确地将Wasserstein距离归因于各类数据成分,包括数据子组、输入特征或可解释子空间。该方法在多种数据集和Wasserstein距离设定下均表现出高准确性,其实际效用在三个应用场景中得到验证。

原文摘要 · Abstract (English)

Wasserstein distances provide a powerful framework for comparing data distributions. They can be used to analyze processes over time or to detect inhomogeneities within data. However, simply calculating the Wasserstein distance or analyzing the corresponding transport plan (or coupling) may not be sufficient for understanding what factors contribute to a high or low Wasserstein distance. In this work, we propose a novel solution based on Explainable AI that allows us to efficiently and accurately attribute Wasserstein distances to various data components, including data subgroups, input features, or interpretable subspaces. Our method achieves high accuracy across diverse datasets and Wasserstein distance specifications, and its practical utility is demonstrated in three use cases.

可解释AI分布对比数据偏移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。