arXiv:2501.07602cs.LGstat.ML2025-01

提出可解释的函数数据机器学习流程,提升高风险场景预测可信度。

An Explainable Pipeline for Machine Learning with Functional Data

  • 结合弹性主成分分析与特征重要性,捕捉函数数据的纵向横向变化
  • 通过可视化重要主成分,揭示模型如何利用数据变异进行预测
  • 适用于材料分类与打印机溯源等高风险领域,强调结果可解释性

机器学习模型在预测任务中表现优异,但其算法复杂性使其难以解释。现有方法多聚焦于一般数据,缺乏针对函数型输入的监督学习可解释性研究。本文针对两个高后果应用场景:基于高光谱断层扫描图像识别爆炸物材料类型,以及利用拉曼光谱提取的色谱签名进行喷墨打印文档源打印机匹配。尽管数据驱动的分类模型是自然选择,但鉴于应用高风险特性,必须合理建模数据本质以避免模式掩盖或误判。为此,本文提出可解释的变量重要性弹性形状分析(VEESA)流程,该流程(1)同时处理函数数据的纵向和横向变异性,(2)在原始数据空间中解释模型如何利用数据变异进行预测。该流程采用弹性函数主成分分析(efPCA)生成不相关的模型输入,并通过置换特征重要性(PFI)识别对预测关键的主成分,其捕获的变异性在原始数据空间中被可视化。最后讨论了流程的自然扩展方向及未来研究挑战。

原文摘要 · Abstract (English)

Machine learning (ML) models have shown success in applications with an objective of prediction, but the algorithmic complexity of some models makes them difficult to interpret. Methods have been proposed to provide insight into these "black-box" models, but there is little research that focuses on supervised ML when the model inputs are functional data. In this work, we consider two applications from high-consequence spaces with objectives of making predictions using functional data inputs. One application aims to classify material types to identify explosive materials given hyperspectral computed tomography scans of the materials. The other application considers the forensics science task of connecting an inkjet printed document to the source printer using color signatures extracted by Raman spectroscopy. An instinctive route to consider for analyzing these data is a data driven ML model for classification, but due to the high consequence nature of the applications, we argue it is important to appropriately account for the nature of the data in the analysis to not obscure or misrepresent patterns. As such, we propose the Variable importance Explainable Elastic Shape Analysis (VEESA) pipeline for training ML models with functional data that (1) accounts for the vertical and horizontal variability in the functional data and (2) provides an explanation in the original data space of how the model uses variability in the functional data for prediction. The pipeline makes use of elastic functional principal components analysis (efPCA) to generate uncorrelated model inputs and permutation feature importance (PFI) to identify the principal components important for prediction. The variability captured by the important principal components in visualized the original data space. We ultimately discuss ideas for natural extensions of the VEESA pipeline and challenges for future research.

可解释AI函数数据高风险预测弹性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。