通过分阶段训练提升推荐系统冷启动表现,避免过度依赖历史曝光数据。
Representation Curriculum: Stagewise Training for Robust Ranking and Allocation

- 先用内容特征训练,再逐步引入曝光信号,防止模型走捷径。
- 在冷启动场景下显著降低风险,且保持头部推荐性能可控下降。
- 适用于电商、推荐等需要公平与鲁棒性的排序系统。
数字市场中的排序是动态的曝光-分配机制:展示内容影响用户发现路径,平台记录的成功事件又用于更新未来策略。现代排序系统高度依赖暴露相关信号(如热度估计、点击率/转化率聚合、ID表示),因其在静态需求下预测力强。然而这种高预测性可能成为学习捷径:早期接触依赖曝光的信念信号会引导优化过度依赖此类信号,忽视内容本身竞争力和语义亲和度等独立于曝光的优质信号。结果导致政策固化既得利益者,恶化冷启动泛化能力,在分布漂移下鲁棒性下降。我们提出表示课程(Representation Curriculum, RC),一种训练时的阶段性特征使用干预方法。RC初期突出内容特征信号,随后逐步引入暴露相关信念信号,同时将内容路径锚定在已学习的优质表示附近,抑制对历史信号的依赖,并缓解内容信号的梯度消失问题。我们在高斯线性岭回归设置中推导出闭式解和充分条件,证明在冷启动目标分布下RC可严格降低总体风险,且量化了与源性能之间的帕累托权衡。在公开学习排序与推荐基准及大规模电商平台的随机在线实验中,RC均显著减少对历史信念信号的依赖,提升冷启动群体表现,且头部性能代价可控。
原文摘要 · Abstract (English)
Ranking in digital marketplaces is a dynamic exposure-allocation mechanism: displayed items shape discovery trajectories and success events logged by the platform to update future allocation policies. Modern ranking systems rely heavily on exposure-confounded signals (e.g. popularity estimates, CTR/CVR aggregates, and ID-based representation), because they are highly predictive under stationary demand. Yet this predictive power can become a learning shortcut: early access to exposure-dependent belief signals steers optimization toward over-reliance on them and away from exposure-independent merit signals (e.g., content-based competitiveness and semantic affinity). Consequently, the learned policy tends to entrench incumbents and degrade cold-start generalization and robustness under distribution shift. We propose Representation Curriculum (RC), a training-time intervention that temporally stages feature utilization. RC foregrounds content-based merit signals initially, then introduces exposure-dependent belief signals while anchoring the content pathway near the learned merit representation, curbing shortcut reliance on historical signals and mitigating gradient starvation on content signals. We formalize RC independently of task and hypothesis class and provide ranking-specific instantiations. In a Gaussian linear ridge setting, we derive closed-form solutions and sufficient conditions under which RC strictly reduces population risk on a cold-start target distribution, with a quantified Pareto tradeoff against source performance. Experiments on public learning-to-rank and recommendation benchmarks, and randomized online experiments in a large-scale e-commerce search system, show that RC measurably shifts reliance from historical belief signals toward content-based merit signals and yields consistent gains on cold populations with a controlled trade-off in head performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。