构建首个高保真多模态时序预测基准,解决评估失真问题。
Fidel-TS: A High-Fidelity Multimodal Benchmark for Time Series Forecasting
- 基于数据溯源完整、无泄露设计的高保真原则构建新基准
- 揭示现有基准评估结果普遍偏高,模型性能存在夸大风险
- 适合时序预测、多模态建模与大模型评估的研究者使用
时序预测模型的评估因缺乏高质量基准而受到阻碍,导致进展评估被高估。现有数据集存在规模小、频率低、单模态设计中预训练数据污染,以及早期多模态设计中常见的时间和描述泄露等问题。为此,我们明确了高保真基准的核心原则:数据来源完整性、无泄露设计与结构清晰性,并据此构建了Fidel-TS这一大规模新基准。实验表明,先前基准存在局限性,模型评估结果可能存在显著偏差,为多种单模态与多模态预测模型及大语言模型在不同任务下的表现提供了新洞见。
原文摘要 · Abstract (English)
The evaluation of time series forecasting models is hindered by a lack of high-quality benchmarks, leading to overestimated assessments of progress. Existing datasets suffer from issues ranging from small-scale, low-frequency, pre-training data contamination in unimodal designs to the temporal and description leakage prevalent in early multimodal designs. To address this, we formalize the core principles of high-fidelity benchmarking, focusing on data sourcing integrity, leak-free design, and structural clarity. We introduce Fidel-TS, a new large-scale benchmark built from these principles. Our experiments reveal the limitations of prior benchmarks and the potential discrepancies in model evaluation, providing new insights into multiple existing unimodal and multimodal forecasting models and LLMs across various evaluation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。