arXiv:2507.11895stat.MLcs.LG2025-07被引 3

提出Newfluence方法,提升高维模型可解释性。

Newfluence: Boosting Model interpretability and Understanding in High Dimensions

  • 基于高维统计理论改进影响函数,解决传统方法失效问题。
  • 在高维场景下显著提升可解释性分析的准确性。
  • 适用于复杂模型诊断,适合研究者与工程师使用。

机器学习与人工智能模型日益复杂,亟需工具帮助科学家、工程师和政策制定者理解并优化模型决策。影响函数源于稳健统计,是当前主流的可解释性方法之一。然而,其启发式基础依赖于参数数p远小于样本数n的低维假设。而现代AI模型常处于p较大的高维区域,挑战这一前提。本文通过理论与实证分析发现,传统影响函数在高维设置中无法可靠工作。为此,我们提出一种新近似方法Newfluence,保持相似计算效率的同时显著提升准确性。Newfluence有望为复杂AI模型提供更可靠的解释与问题诊断能力。此外,本文构建的高维框架还可用于分析其他方法(如Shapley值)。

原文摘要 · Abstract (English)

The increasing complexity of machine learning (ML) and artificial intelligence (AI) models has created a pressing need for tools that help scientists, engineers, and policymakers interpret and refine model decisions and predictions. Influence functions, originating from robust statistics, have emerged as a popular approach for this purpose. However, the heuristic foundations of influence functions rely on low-dimensional assumptions where the number of parameters $p$ is much smaller than the number of observations $n$. In contrast, modern AI models often operate in high-dimensional regimes with large $p$, challenging these assumptions. In this paper, we examine the accuracy of influence functions in high-dimensional settings. Our theoretical and empirical analyses reveal that influence functions cannot reliably fulfill their intended purpose. We then introduce an alternative approximation, called Newfluence, that maintains similar computational efficiency while offering significantly improved accuracy. Newfluence is expected to provide more accurate insights than many existing methods for interpreting complex AI models and diagnosing their issues. Moreover, the high-dimensional framework we develop in this paper can also be applied to analyze other popular techniques, such as Shapley values.

可解释性高维数据模型诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。