通过辅助学习提升推荐系统对少数用户群体的捕捉能力
Improving Large-Scale Recommender Systems with Auxiliary Learning
- 用部分冲突的辅助标签正则化共享表示,增强注意力机制
- 在百亿数据上实验,使少数群体性能提升超0.30%
- 适合大规模推荐系统中关注长尾用户的研究者
在单一全局目标下训练大规模推荐模型隐含假设用户群体同质性,但真实数据由具有不同条件分布的异构群体构成。随着模型规模与数据量增大,模型逐渐被中心分布主导,忽略头部与尾部区域,导致学习能力受限,出现注意力权重失效或神经元死亡。本文揭示注意力机制在因子分解机中共享嵌入选择中的关键作用,提出通过分析数据子结构并利用辅助学习暴露具有强分布差异的群体来解决该问题。不同于以往通过加权标签或多任务头启发式缓解偏差的方法,本方法使用部分冲突的辅助标签对共享表示进行正则化,定制注意力层学习过程,在保留与少数群体互信息的同时提升全局性能。在包含数十亿数据点的六个SOTA模型的生产级数据集上评估,结果显示因子分解机能更精细捕捉用户-广告交互,整体归一化熵降低最高达0.16%,对目标少数群体性能提升超过0.30%。
原文摘要 · Abstract (English)
Training large-scale recommendation models under a single global objective implicitly assumes homogeneity across user populations. However, real-world data are composites of heterogeneous cohorts with distinct conditional distributions. As models increase in scale and complexity and as more data is used for training, they become dominated by central distribution patterns, neglecting head and tail regions. This imbalance limits the model's learning ability and can result in inactive attention weights or dead neurons. In this paper, we reveal how the attention mechanism can play a key role in factorization machines for shared embedding selection, and propose to address this challenge by analyzing the substructures in the dataset and exposing those with strong distributional contrast through auxiliary learning. Unlike previous research, which heuristically applies weighted labels or multi-task heads to mitigate such biases, we leverage partially conflicting auxiliary labels to regularize the shared representation. This approach customizes the learning process of attention layers to preserve mutual information with minority cohorts while improving global performance. We evaluated proposed method on massive production datasets with billions of data points each for six SOTA models. Experiments show that the factorization machine is able to capture fine-grained user-ad interactions using the proposed method, achieving up to a 0.16% reduction in normalized entropy overall and delivering gains exceeding 0.30% on targeted minority cohorts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。