用两种高效算法组合提升时间序列分类准确率,突破传统模型效率瓶颈
The Meta-Learning Gap: Combining Hydra and Quant for Large-Scale Time Series Classification
- 融合卷积核竞争与分段量化方法,构建轻量级集成模型
- 在10个大规模数据集上平均准确率提升至0.836,7个数据集表现超越基线
- 揭示组合策略不足导致的元学习差距,适合优化集成学习的研究者参考
时间序列分类面临准确率与计算效率之间的根本权衡。尽管综合集成方法HIVE-COTE 2.0在UCR基准上达到最先进准确率,但其340小时的训练时间使其难以应用于大规模数据集。本文研究是否可通过互补范式中两种高效算法的针对性组合,兼顾集成优势与计算可行性。将针对性卷积核竞争的Hydra与分段区间量化方法Quant在六个集成配置下结合,评估其在10个大规模MONSTER数据集(含7,898至1,168,774个训练样本)上的表现。最强配置使平均准确率从0.829提升至0.836,在10个数据集中有7个取得成功。然而,预测组合型集成仅捕获了理论最优解的11%,揭示出显著的元学习优化差距。特征拼接方法通过学习新决策边界超出最优解表现,而预测层面的互补性与集成增益呈中等相关。核心发现:挑战已从确保算法差异转向如何有效组合;当前元学习策略难以利用存在互补性的事实。改进组合策略或可使多种时间序列分类应用中的集成增益翻倍甚至三倍。
原文摘要 · Abstract (English)
Time series classification faces a fundamental trade-off between accuracy and computational efficiency. While comprehensive ensembles like HIVE-COTE 2.0 achieve state-of-the-art accuracy, their 340-hour training time on the UCR benchmark renders them impractical for large-scale datasets. We investigate whether targeted combinations of two efficient algorithms from complementary paradigms can capture ensemble benefits while maintaining computational feasibility. Combining Hydra (competing convolutional kernels) and Quant (hierarchical interval quantiles) across six ensemble configurations, we evaluate performance on 10 large-scale MONSTER datasets (7,898 to 1,168,774 training instances). Our strongest configuration improves mean accuracy from 0.829 to 0.836, succeeding on 7 of 10 datasets. However, prediction-combination ensembles capture only 11% of theoretical oracle potential, revealing a substantial meta-learning optimization gap. Feature-concatenation approaches exceeded oracle bounds by learning novel decision boundaries, while prediction-level complementarity shows moderate correlation with ensemble gains. The central finding: the challenge has shifted from ensuring algorithms are different to learning how to combine them effectively. Current meta-learning strategies struggle to exploit the complementarity that oracle analysis confirms exists. Improved combination strategies could potentially double or triple ensemble gains across diverse time series classification applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。