arXiv:2504.18329cs.LGcs.AI2025-04被引 1

用拓扑方法自动筛选时间序列关键特征,既省计算又易解释。

PHEATPRUNER: Interpretable Data-centric Feature Selection for Multivariate Time Series Classification through Persistent Homology

  • 结合持久同调与层论,从多变量时序数据中提取结构特征
  • 可剪枝高达45%变量,模型准确率不降反升
  • 适合需要透明决策过程的医疗、工业监测等场景

多变量时间序列分类中,性能与可解释性难以兼顾,源于数据复杂性和高维性。本文提出PHeatPruner,融合持久同调与层论,实现数据驱动的特征选择。该方法可在不依赖后验概率或监督优化算法的前提下,对随机森林、CatBoost、XGBoost和LightGBM等模型剪枝高达45%的输入变量,同时保持或提升模型准确率。层论提供解释向量,揭示数据深层结构。在UEA Archive和奶牛乳腺炎检测数据集上验证有效,显著简化复杂数据,且处理时间与复杂度未增加。该方法打通了降维与可解释性之间的鸿沟,适用于多个实际领域。

原文摘要 · Abstract (English)

Balancing performance and interpretability in multivariate time series classification is a significant challenge due to data complexity and high dimensionality. This paper introduces PHeatPruner, a method integrating persistent homology and sheaf theory to address these challenges. Persistent homology facilitates the pruning of up to 45% of the applied variables while maintaining or enhancing the accuracy of models such as Random Forest, CatBoost, XGBoost, and LightGBM, all without depending on posterior probabilities or supervised optimization algorithms. Concurrently, sheaf theory contributes explanatory vectors that provide deeper insights into the data's structural nuances. The approach was validated using the UEA Archive and a mastitis detection dataset for dairy cows. The results demonstrate that PHeatPruner effectively preserves model accuracy. Furthermore, our results highlight PHeatPruner's key features, i.e. simplifying complex data and offering actionable insights without increasing processing time or complexity. This method bridges the gap between complexity reduction and interpretability, suggesting promising applications in various fields.

时间序列特征选择拓扑分析可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。