arXiv:2510.25108cs.LGstat.ML2025-10被引 5

训练时混入不匹配数据,反而能提升测试表现。

Shift is Good: Mismatched Data Mixing Improves Test Performance

  • 通过调整训练数据比例制造分布偏移,提升模型泛化能力。
  • 在多种场景下,特定训练比例可使测试性能显著提高。
  • 适用于需要鲁棒性的实际应用,如跨域任务、技能分布不一致场景。

我们研究在训练和测试时使用不同比例混合分布的情况。结果显示,在许多情形下,且在某种意义上具有普遍性,分布偏移反而有益,即使各组件无关且无知识迁移,测试性能仍可能因训练比例不匹配而提升。在多种场景中,我们确定了最优训练比例,并量化了分布偏移带来的增益程度。该分析同样适用于组件‘技能’分布不同的组合式设置中训练与测试的差异。

原文摘要 · Abstract (English)

We consider training and testing on mixture distributions with different training and test proportions. We show that in many settings, and in some sense generically, distribution shift can be beneficial, and test performance can improve due to mismatched training proportions, even if the components are unrelated and with no transfer between components. In a variety of scenarios, we identify the optimal training proportions and the extent to which such distribution shift can be beneficial. We show how the same analysis applies also to a compositional setting with differing distribution of component "skills'' at training and test.

分布偏移训练策略泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。