预训练数据分布影响模型少样本学习能力,存在鲁棒性与泛化性的根本权衡。
How Does the Pretraining Distribution Shape In-Context Learning? A Fundamental Trade-Off
- 通过理论框架分析预训练分布特性对少样本学习的影响
- 重尾分布提升抗分布偏移能力,但降低低数据下的泛化性能
- 在随机微分方程等复杂任务上验证了预测的可靠性
尽管少样本学习(ICL)在大型语言模型中表现出惊人效果,能仅凭少量示例适应新任务,但其性能驱动因素仍不明确。本文研究预训练数据分布的统计特性(如尾部行为、覆盖范围)如何塑造ICL能力。我们构建了一个涵盖泛化与任务选择的理论框架,揭示分布特性如何决定样本效率、任务检索和鲁棒性。通过将现有集中度结果推广至重尾先验和依赖序列,更准确反映大模型预训练数据结构。研究发现存在根本性设计权衡:重尾分布有助于在分布偏移下实现稳健的任务选择,却损害低数据场景下的泛化能力。我们在随机微分方程及带记忆的随机过程等挑战性任务上实证验证了这些预测。结果表明,控制预训练分布的关键统计特性,对构建具备稳定少样本学习能力的模型至关重要。
原文摘要 · Abstract (English)
The factors driving the performance of in-context learning (ICL) in large language models (LLMs) remain poorly understood despite ICL's surprising effectiveness, enabling models to adapt to new tasks from only a handful of examples. To clarify and improve these capabilities, we characterize how the statistical properties of the pretraining distribution (e.g., tail behavior, coverage) shape ICL. We develop a theoretical framework that encompasses generalization and task selection and show how distributional properties govern sample efficiency, task retrieval, and robustness. To this end, we generalize existing concentration results to heavy-tailed priors and dependent sequences, better reflecting the structure of LLM pretraining data. Our framework reveals a fundamental design trade-off: heavy-tailed pretraining distributions facilitate robust task selection under distribution shifts but are detrimental to generalization, especially in low-data regimes. We then empirically evaluate our predictions by studying how ICL performance varies with the pretraining distribution on challenging tasks such as stochastic differential equations and stochastic processes with memory. Together, these findings suggest that controlling key statistical properties of the pretraining distribution is essential for building ICL-capable and reliable LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。