arXiv:2511.06374cs.LGstat.ML2025-11被引 1

针对大规模稀疏特征模型多轮训练过拟合问题,提出自适应正则化方法提升性能。

Adaptive Regularization for Large-Scale Sparse Feature Embedding Models

  • 基于Rademacher复杂度理论,分析稀疏特征过拟合成因
  • 自适应约束嵌入层范数预算,单轮训练性能提升,多轮无下降
  • 已落地线上系统,适用于搜索推荐等场景的高维稀疏模型

单轮训练过拟合问题在搜索、广告和推荐领域的点击率(CTR)与转化率(CVR)预测模型中备受关注。这些依赖大规模稀疏类别特征的模型在多轮训练时性能显著下降。尽管已有研究提出启发式解决方案,但根本原因仍不明确。本文基于Rademacher复杂度提供理论解释,并通过实验证实。据此提出一种自适应正则化方法,动态约束嵌入层的范数预算。该方法不仅避免了多轮训练中的严重性能退化,还提升了单轮训练表现。该方法已在生产系统中部署。

原文摘要 · Abstract (English)

The one-epoch overfitting problem has drawn widespread attention, especially in CTR and CVR estimation models in search, advertising, and recommendation domains. These models which rely heavily on large-scale sparse categorical features, often suffer a significant decline in performance when trained for multiple epochs. Although recent studies have proposed heuristic solutions, the fundamental cause of this phenomenon remains unclear. In this work, we present a theoretical explanation grounded in Rademacher complexity, supported by empirical experiments, to explain why overfitting occurs in models with large-scale sparse categorical features. Based on this analysis, we propose a regularization method that constrains the norm budget of embedding layers adaptively. Our approach not only prevents the severe performance degradation observed during multi-epoch training, but also improves model performance within a single epoch. This method has already been deployed in online production systems.

稀疏特征正则化在线学习深度推荐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。