arXiv:2507.13287stat.MEcs.LG2025-07被引 1

提出RIDER方法,有效应对数据分布随时间变化的挑战。

Optimal Empirical Risk Minimization under Temporal Distribution Shifts

  • 基于随机分布漂移模型,推导出最优加权学习策略
  • 在多个真实数据集上显著提升预测性能
  • 适合动态环境中的模型持续优化,如金融与交通预测

时间分布漂移是动态环境中机器学习模型面临的核心挑战。本文提出RIDER(RIsk minimization under Dynamically Evolving Regimes),在时间分布漂移下推导出最优加权的经验风险最小化方法。该方法基于随机分布漂移模型,其中随机漂移源于数据生成过程的大量不可预测变化的叠加。我们证明,常见的加权策略——如合并全部数据、指数加权数据、仅使用最新数据——均自然地成为本框架的特例。实验表明,将RIDER作为微调步骤应用于Yearbook数据集,在Wild-Time基准方法中持续提升样本外预测性能。此外,RIDER在两个真实任务中也优于标准加权策略:预测股票市场波动率和纽约出租车行程时长预测。

原文摘要 · Abstract (English)

Temporal distribution shifts pose a key challenge for machine learning models trained and deployed in dynamically evolving environments. This paper introduces RIDER (RIsk minimization under Dynamically Evolving Regimes) which derives optimally-weighted empirical risk minimization procedures under temporal distribution shifts. Our approach is theoretically grounded in the random distribution shift model, where random shifts arise as a superposition of numerous unpredictable changes in the data-generating process. We show that common weighting schemes, such as pooling all data, exponentially weighting data, and using only the most recent data, emerge naturally as special cases in our framework. We demonstrate that RIDER consistently improves out-of-sample predictive performance when applied as a fine-tuning step on the Yearbook dataset, across a range of benchmark methods in Wild-Time. Moreover, we show that RIDER outperforms standard weighting strategies in two other real-world tasks: predicting stock market volatility and forecasting ride durations in NYC taxi data.

分布漂移时间序列模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。