用超像素思想加速大规模时间序列的可解释性分析,用于法庭级DNA分类。
A novel application of Shapley values for large multidimensional time-series data: Applying explainable AI to a DNA profile classification neural network
- 将超像素思路引入Shapley值计算,降低高维时间序列分析复杂度
- 在31,200个扫描点的DNA数据上实现快速、准确的解释性分析
- 为司法鉴定中的人工读图提供自动化替代方案,适合法医与AI交叉研究者
将Shapley值应用于高维时序数据存在计算挑战——对于N个输入,问题复杂度为2^N。在图像处理中,通过将像素聚类为超像素来简化计算。本研究提出一种适用于时序类数据的高效方法,借鉴超像素思想优化Shapley值计算。以法医DNA分类为例,该方法应用于由卷积神经网络(CNN)处理的多变量时序数据。在DNA分析中,需从提取和处理产生的背景噪声中识别等位基因,单个DNA谱图包含31,200个扫描点,分类结果必须在法庭上可辩护,因此长期依赖人工判读,耗时巨大。本研究展示了一种可实现快速、真实且准确的Shapley值计算,为该任务提供了潜在的自动化替代方案。
原文摘要 · Abstract (English)
The application of Shapley values to high-dimensional, time-series-like data is computationally challenging - and sometimes impossible. For $N$ inputs the problem is $2^N$ hard. In image processing, clusters of pixels, referred to as superpixels, are used to streamline computations. This research presents an efficient solution for time-seres-like data that adapts the idea of superpixels for Shapley value computation. Motivated by a forensic DNA classification example, the method is applied to multivariate time-series-like data whose features have been classified by a convolutional neural network (CNN). In DNA processing, it is important to identify alleles from the background noise created by DNA extraction and processing. A single DNA profile has $31,200$ scan points to classify, and the classification decisions must be defensible in a court of law. This means that classification is routinely performed by human readers - a monumental and time consuming process. The application of a CNN with fast computation of meaningful Shapley values provides a potential alternative to the classification. This research demonstrates the realistic, accurate and fast computation of Shapley values for this massive task
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。