动态任务分布提升模型泛化能力,减少记忆倾向。
Temporal Task Diversity: Inductive Biases Under Non-Stationarity in Synthetic Sequence Modelling

- 在训练中引入随时间变化的任务多样性
- 模型更倾向于泛化而非记忆训练数据
- 适合研究非平稳环境下的学习机制
现代深度学习常假设神经网络从固定数据分布中学习,但许多实际问题涉及随训练过程变化的数据分布。这种非平稳性如何影响深度学习的归纳偏置,即对不同结构、泛化和安全性模型的偏好?一个有效的研究归纳偏置的测试平台是上下文内线性回归序列建模,其中小型Transformer在固定任务分布下表现出显著不同的泛化模式。本文探索了在训练过程中任务分布的多样化效应,发现这种时间上的多样性会增强模型向泛化而非记忆的偏倚。
原文摘要 · Abstract (English)
Modern deep learning science often assumes that neural networks learn from a fixed data distribution. However, many practically important learning problems involve data distributions that change throughout training. How does such non-stationarity impact the inductive biases of deep learning towards models with different structural, generalisation, and safety properties? A fruitful testbed for studying inductive bias is in-context linear regression sequence modelling, where small transformers display strikingly different generalisation patterns depending on the diversity of the (fixed) training task distribution. In this paper, we explore the effect of diversifying the task distribution across training time, finding that such temporal diversity leads to an increased bias towards generalisation over memorisation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。