arXiv:2412.01460cs.DBcs.LG2024-12中稿 · VLDB 2025被引 5

系统梳理了沙普利值在数据分析中的应用与挑战

A Comprehensive Study of Shapley Value in Data Analytics

  • 提出四类核心挑战:计算效率、近似误差、隐私保护和可解释性
  • 构建SVBench框架,验证不同技术在各类任务中的表现
  • 适合关注可解释性与数据科学工具链的研究者

近年来,合作博弈论中的沙普利值(Shapley Value, SV)在数据分析(DA)领域得到广泛应用。本文首次全面研究了SV在整个数据分析流程中的应用,明确了适用于数据分析的SV定义关键变量及其为数据科学家提供的核心功能。我们总结了使用SV面临的四大挑战:计算效率、近似误差、隐私保护和可解释性,厘清现有方法在应对这些挑战时的技术路径,并分析其在不同挑战间的潜在冲突。此外,我们开发了SVBench——一个模块化、可扩展的开源框架,用于在各类数据分析任务中构建SV应用,并通过大量实验验证了分析结论。基于定性和定量结果,本文揭示了当前应用的局限性,指出了未来研究与工程的方向。

原文摘要 · Abstract (English)

Over the recent years, Shapley value (SV), a solution concept from cooperative game theory, has found numerous applications in data analytics (DA). This paper presents the first comprehensive study of SV used throughout the DA workflow, clarifying the key variables in defining DA-applicable SV and the essential functionalities that SV can provide for data scientists. We condense four primary challenges of using SV in DA, namely computation efficiency, approximation error, privacy preservation, and interpretability, disentangle the resolution techniques from existing arts in this field, then analyze and discuss the techniques w.r.t. each challenge and the potential conflicts between challenges.We also implement SVBench, a modular and extensible open-source framework for developing SV applications in different DA tasks, and conduct extensive evaluations to validate our analyses and discussions. Based on the qualitative and quantitative results, we identify the limitations of current efforts for applying SV to DA and highlight the directions of future research and engineering.

沙普利值可解释性数据分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。