用数据指纹预测算法性能,无需训练即可选最优模型
Utilizing Data Fingerprints for Privacy-Preserving Algorithm Selection in Time Series Classification: Performance and Uncertainty Estimation on Unseen Datasets
- 构建时间序列数据指纹,隐私保护下表征数据特征
- 在112个数据集上预测35种算法表现,平均误差降低7.32%
- 适合需快速选型且数据敏感的工业级时序分类场景
时间序列分类中算法选择至关重要。传统方法如神经架构搜索、自动化机器学习等虽有效,但需大量计算资源并依赖完整数据训练。本文提出一种新型数据指纹,以隐私保护方式表征任意时间序列分类数据集,并在不训练(未见)数据集的情况下提供算法选择洞察。通过分解多目标回归问题,仅使用数据指纹即可可扩展、自适应地估计算法性能与不确定性。在112个加州大学河滨基准数据集上评估,成功预测35种先进算法表现,相比朴素基线平均提升7.32%的均值性能预测准确率和15.81%的不确定性估计精度。
原文摘要 · Abstract (English)
The selection of algorithms is a crucial step in designing AI services for real-world time series classification use cases. Traditional methods such as neural architecture search, automated machine learning, combined algorithm selection, and hyperparameter optimizations are effective but require considerable computational resources and necessitate access to all data points to run their optimizations. In this work, we introduce a novel data fingerprint that describes any time series classification dataset in a privacy-preserving manner and provides insight into the algorithm selection problem without requiring training on the (unseen) dataset. By decomposing the multi-target regression problem, only our data fingerprints are used to estimate algorithm performance and uncertainty in a scalable and adaptable manner. Our approach is evaluated on the 112 University of California riverside benchmark datasets, demonstrating its effectiveness in predicting the performance of 35 state-of-the-art algorithms and providing valuable insights for effective algorithm selection in time series classification service systems, improving a naive baseline by 7.32% on average in estimating the mean performance and 15.81% in estimating the uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。