构建金融时序预测新基准,让模型评估更贴近真实投资决策。
FinVerse: Financial Time-Series Benchmark

- 按经济意义匹配11类78项指标,取代统一误差评估。
- 覆盖11.6万条金融序列,其中1.7万条为高经济相关性目标序列。
- 揭示通用模型在标准指标下表现好,但未必适合实际金融决策。
随着时间序列基础模型的兴起,评估其预测能力的基准变得愈发重要。现有时间序列预测基准虽提供标准化对比,但通常对异质序列使用统一的误差指标,强表现未必意味着支持最优实际决策。例如,在股票预测中,准确判断涨跌比最小化点预测误差更具现实意义。为此,我们推出FinVerse——首个面向金融领域的时序预测基准,迈出更真实评估的第一步。该数据集包含116,897条金融时序,共1711万条观测,其中60,232条(1740万条观测)根据其与金融决策的经济相关性被选为评估目标。不同于侧重统一点预测或概率精度的通用基准,FinVerse基于每条序列的经济含义,为其分配最合适的评估指标,涵盖11类共78项指标。对43个公开时间序列基础模型的分析显示,通用指标下的优异表现,并不等同于产生有用的金融预测。这凸显了面向领域、贴近真实决策目标的评估基准的必要性。
原文摘要 · Abstract (English)
As time-series foundation models have emerged, the need for benchmarks that can evaluate their forecasting ability in meaningful ways has become increasingly important. Existing time-series forecasting benchmarks provide useful standardized comparisons, but they often evaluate heterogeneous series with uniform error-based metrics. Strong performance under such metrics does not necessarily imply that a model's forecasts will support the best real-world decisions across domains. For example, in stock forecasting, correctly predicting whether a price will rise or fall can be more directly relevant to realized returns than minimizing point-wise forecast error alone. To this end, we introduce FinVerse, a finance-domain time-series forecasting benchmark that takes a first step toward more realistic evaluation. The released FinVerse data artifact contains 116,897 financial time series with 171.1M observations, of which 60,232 series with 17.4M observations are selected as evaluated targets based on their economic relevance to financial decisions. Unlike generic forecasting benchmarks that primarily emphasize uniform point-forecast or probabilistic accuracy, FinVerse defines 11 metric families comprising 78 evaluation metrics and assigns the most appropriate evaluation metrics to each individual time series based on its underlying economic meaning. Our analysis of 43 public time-series forecasting foundation models shows that strong performance under generic forecasting criteria does not necessarily translate into useful financial forecasts. This finding highlights the need for domain-aware benchmarks that evaluate models under objectives closer to real-world decision making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。