arXiv:2603.22074cs.LG2026-03被引 1

用多实例学习提升时间序列分类准确率与可解释性

MIHT: A Hoeffding Tree for Time Series Classification using Multiple Instance Learning

  • 将时序数据视为子序列集合,通过增量决策树识别关键片段
  • 在28个公开数据集上超越11种先进模型,高维数据表现优异
  • 结果可解释,适合需要理解判别依据的场景

由于时间序列在现实问题中普遍存在且具有内在依赖性,时间序列分类在多个领域至关重要。然而,现有模型常难以处理变长或高维序列。本文提出MIHT(多实例霍夫丁树)算法,一种高效模型,利用多实例学习对多变量、变长时间序列进行分类,并提供可解释结果。该方法将时间序列表示为“子序列袋”,结合基于增量决策树的优化过程,区分序列中的相关部分与噪声。该方法提取多变量、变长序列的潜在概念。生成的决策树是序列概念的紧凑白盒表示,可揭示最相关的变量和片段。实验表明,MIHT在28个公共数据集上优于11种先进时序分类模型,包括高维数据集。该模型兼具更高准确率与可解释性,是处理复杂动态时序数据的有力方案。

原文摘要 · Abstract (English)

Due to the prevalence of temporal data and its inherent dependencies in many real-world problems, time series classification is of paramount importance in various domains. However, existing models often struggle with series of variable length or high dimensionality. This paper introduces the MIHT (Multi-instance Hoeffding Tree) algorithm, an efficient model that uses multi-instance learning to classify multivariate and variable-length time series while providing interpretable results. The algorithm uses a novel representation of time series as "bags of subseries," together with an optimization process based on incremental decision trees that distinguish relevant parts of the series from noise. This methodology extracts the underlying concept of series with multiple variables and variable lengths. The generated decision tree is a compact, white-box representation of the series' concept, providing interpretability insights into the most relevant variables and segments of the series. Experimental results demonstrate MIHT's superiority, as it outperforms 11 state-of-the-art time series classification models on 28 public datasets, including high-dimensional ones. MIHT offers enhanced accuracy and interpretability, making it a promising solution for handling complex, dynamic time series data.

时间序列分类多实例学习可解释性决策树

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。