arXiv:2510.16322cs.LGstat.ML2025-10

记忆长尾数据能提升模型对罕见组合的泛化能力

Memorizing Long-tail Data Can Help Generalization Through Composition

  • 通过记忆长尾特征,结合组合推理机制提升泛化
  • 在未见过的罕见特征组合上仍能正确预测
  • 模型架构影响组合能力,适合研究长尾分布场景

深度学习促使人们重新思考记忆与泛化之间的关系。在许多场景中,记忆不会损害泛化能力,反而可能通过记住长尾样本而带来帮助。本文探讨记忆与简单组合能力之间的协同作用——即模型能否对长尾特征的组合做出正确预测。理论上,我们证明在线性设定下,记忆与组合能力相结合,可使模型在测试时对从未在训练中出现过的长尾特征组合做出正确预测。神经网络实验表明,这一理论洞察可扩展至非线性情形,且模型的组合能力与其架构密切相关。

原文摘要 · Abstract (English)

Deep learning has led researchers to rethink the relationship between memorization and generalization. In many settings, memorization does not hurt generalization due to implicit regularization and may help by memorizing long-tailed examples. In this paper, we consider the synergy between memorization and simple composition -- the ability to make correct prediction on a combination of long-tailed features. Theoretically, we show that for a linear setting, memorization together with composition can help the model make correct predictions on rare test examples that require a combination of long-tailed features, even if such combinations were never observed in the training data. Experiments on neural network architecture on simple data show that the theoretical insight extends beyond the linear setting, and we further observe that the composition capability of the model depends on its architecture.

记忆机制长尾分布组合泛化模型架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。