首个综合评估生成时间序列的基准框架,可自动对比41种评估方法。
STEB: In Search of the Best Evaluation Approach for Synthetic Time Series
- 构建多数据集、可配置变换的自动化评估框架
- 验证41种评估指标并发现嵌入方式影响评分结果
- 适合研究生成模型评估或数据隐私应用的开发者
由于数据增强或隐私法规需求,合成时间序列日益重要,催生了众多生成模型、框架和评估方法。然而,大规模客观比较这些评估方法仍具挑战。本文提出合成时间序列评估基准(STEB),首个支持全面、可解释的自动化比较框架。基于10个不同数据集,通过随机注入与13种可配置变换,STEB计算评估指标的可靠性与得分一致性,并记录运行时间、测试误差,支持串行与并行模式。实验中对41种文献中的评估方法进行排序,确认上游时间序列嵌入方式显著影响最终评分。
原文摘要 · Abstract (English)
The growing need for synthetic time series, due to data augmentation or privacy regulations, has led to numerous generative models, frameworks, and evaluation measures alike. Objectively comparing these measures on a large scale remains an open challenge. We propose the Synthetic Time series Evaluation Benchmark (STEB) -- the first benchmark framework that enables comprehensive and interpretable automated comparisons of synthetic time series evaluation measures. Using 10 diverse datasets, randomness injection, and 13 configurable data transformations, STEB computes indicators for measure reliability and score consistency. It tracks running time, test errors, and features sequential and parallel modes of operation. In our experiments, we determine a ranking of 41 measures from literature and confirm that the choice of upstream time series embedding heavily impacts the final score.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。