arXiv:2601.12296cs.LG2026-01AAAI

分布差异越大,模型越能学到稳定预测,即使只用普通训练方法。

Distribution Shift Is Key to Learning Invariant Prediction

  • 通过分析训练数据的分布差异,发现其直接影响模型预测能力。
  • 分布差异大时,普通方法也能逼近最优的不变预测模型。
  • 适合关注模型泛化性与鲁棒性的研究者阅读。

一种有趣现象是:经验风险最小化(ERM)有时反而优于专为分布外任务设计的方法。本研究揭示,这背后的原因之一在于训练域间的分布差异。我们发现,较大的分布差异能显著提升模型性能,甚至使普通方法接近不变预测模型。理论上,我们推导出上界,表明分布差异越大,模型预测能力越强,越接近在任意已知或未知域下保持稳定的不变预测;反之则弱。同时证明,在特定数据条件下,ERM 解可达到与不变预测模型相当的性能。实验验证显示,当训练数据的分布差异增大时,模型预测结果趋近于理想或最优模型。

原文摘要 · Abstract (English)

An interesting phenomenon arises: Empirical Risk Minimization (ERM) sometimes outperforms methods specifically designed for out-of-distribution tasks. This motivates an investigation into the reasons behind such behavior beyond algorithmic design. In this study, we find that one such reason lies in the distribution shift across training domains. A large degree of distribution shift can lead to better performance even under ERM. Specifically, we derive several theoretical and empirical findings demonstrating that distribution shift plays a crucial role in model learning and benefits learning invariant prediction. Firstly, the proposed upper bounds indicate that the degree of distribution shift directly affects the prediction ability of the learned models. If it is large, the models' ability can increase, approximating invariant prediction models that make stable predictions under arbitrary known or unseen domains; and vice versa. We also prove that, under certain data conditions, ERM solutions can achieve performance comparable to that of invariant prediction models. Secondly, the empirical validation results demonstrated that the predictions of learned models approximate those of Oracle or Optimal models, provided that the degree of distribution shift in the training data increases.

分布外泛化不变学习理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。