用连续谢林值解析函数型数据模型的预测贡献
Functional relevance based on the continuous Shapley value
- 基于连续谢林值构建函数型数据的可解释性方法
- 在模拟与真实数据上验证了方法的有效性
- 适合研究函数型数据分析与模型解释的学者
人工智能在社会中的应用日益广泛,理解其行为(如基于表格、文本或图像的机器学习预测算法)变得愈发重要。本文聚焦于函数型数据预测模型的可解释性问题。由于函数型数据特征空间为无限维,传统解释方法难以适用。为此,针对标量对函数回归任务,提出基于连续谢林值的可解释性方法,该数学框架能公平分配无限参与者群体的全局收益。通过模拟与真实数据集实验验证了方法性能,并发布了开源Python工具包ShapleyFDA。
原文摘要 · Abstract (English)
The presence of artificial intelligence (AI) in our society is increasing, which brings with it the need to understand the behavior of AI mechanisms, including machine learning predictive algorithms fed with tabular data, text or images, among others. This work focuses on interpretability of predictive models based on functional data. Designing interpretability methods for functional data models implies working with a set of features whose size is infinite. In the context of scalar on function regression, we propose an interpretability method based on the Shapley value for continuous games, a mathematical formulation that allows for the fair distribution of a global payoff among a continuous set of players. The method is illustrated through a set of experiments with simulated and real data sets. The open source Python package ShapleyFDA is also presented.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。