arXiv:2509.23487cs.LGcs.CL2025-09

检验模型用历史数据预测未来的能力,发现现有方法均不如直接用最新模型。

Temporal Generalization: A Reality Check

  • 通过参数插值和外推两种方式尝试从历史模型预测未来表现。
  • 所有方法在多数任务上都未能超越仅使用最新模型的简单基准。
  • 提醒研究者警惕对时间外推能力的过度宣称,尤其缺乏未来数据时。

机器学习模型在面对分布变化时往往难以保持性能,导致对未来未见数据的预测不准确。本文研究了仅依赖历史数据能否实现对未来的泛化,并探讨了两种主要方法:过去模型参数的凸组合(参数插值)与超出历史参数凸包的显式外推(参数外推)。我们在包括语言建模、新闻摘要、新闻标签预测、论文分类、卫星图像土地利用分类及历史年鉴照片性别预测在内的多种时间任务上进行了评估。实验结果表明,在所有场景中,没有一种方法能持续优于仅使用最新可用模型参数的简单基线。在无法获取未来数据或对数据生成过程无稳健假设的情况下,这些结果凸显了向未来泛化与外推的内在困难,警示人们在评估此类泛化能力时应保持谨慎。

原文摘要 · Abstract (English)

Machine learning (ML) models often struggle to maintain performance under distribution shifts, leading to inaccurate predictions on unseen future data. In this work, we investigate whether and under what conditions models can achieve such a generalization when relying solely on past data. We explore two primary approaches: convex combinations of past model parameters (\emph{parameter interpolation}) and explicit extrapolation beyond the convex hull of past parameters (\emph{parameter extrapolation}). We benchmark several methods within these categories on a diverse set of temporal tasks, including language modeling, news summarization, news tag prediction, academic paper categorization, satellite image-based land use classification over time, and historical yearbook photo gender prediction. Our empirical findings show that none of the evaluated methods consistently outperforms the simple baseline of using the latest available model parameters in all scenarios. In the absence of access to future data or robust assumptions about the underlying data-generating process, these results underscore the inherent difficulties of generalizing and extrapolating to future data and warrant caution when evaluating claims of such generalization.

时间泛化分布外模型外推

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。