arXiv:2409.15582stat.MLcond-mat.dis-nn2024-09NeurIPS被引 2

研究概念漂移下模型泛化与专精的权衡,发现长上下文可能损害预测性能。

Generalization vs. Specialization under Concept Shift

  • 基于岭回归理论推导热力学极限下的预测风险表达式
  • 揭示概念漂移导致测试性能非单调变化,存在弱强漂移相变
  • 实验证明大模型在长上下文时泛化能力反而下降,适用于分类任务

机器学习模型在分布偏移下常表现脆弱,尤其当测试时输入-标签关系发生变化(即概念漂移)时。本文分析了岭回归在概念漂移下的表现,推导出热力学极限下的精确预测风险公式。结果表明概念漂移对泛化性能有非平凡影响:存在弱与强概念漂移之间的相变,且即使不存在双下降现象,测试性能仍呈现非单调的数据依赖性。基于预训练变换器解决线性回归的实验显示,在概念漂移下过长的上下文长度会损害下一词预测的泛化性能。在MNIST和FashionMNIST上的实验进一步表明,这种现象同样存在于分类任务中。

原文摘要 · Abstract (English)

Machine learning models are often brittle under distribution shift, i.e., when data distributions at test time differ from those during training. Understanding this failure mode is central to identifying and mitigating safety risks of mass adoption of machine learning. Here we analyze ridge regression under concept shift -- a form of distribution shift in which the input-label relationship changes at test time. We derive an exact expression for prediction risk in the thermodynamic limit. Our results reveal nontrivial effects of concept shift on generalization performance, including a phase transition between weak and strong concept shift regimes and nonmonotonic data dependence of test performance even when double descent is absent. Our theoretical results are in good agreement with experiments based on transformers pretrained to solve linear regression; under concept shift, too long context length can be detrimental to generalization performance of next token prediction. Finally, our experiments on MNIST and FashionMNIST suggest that this intriguing behavior is present also in classification problems.

概念漂移泛化能力大模型理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。