arXiv:2607.23749cs.IR2026-07中稿 · 20th ACM Conferenc…

破解音乐推荐中的循环困局,提升新歌发现与多样性。

Breaking the Loop: An Empirical Comparison of Strategies for Novelty and Freshness in YouTube Music

论文配图:Breaking the Loop: An Empirical Comparison of Strategies for Novelty and Freshness in YouTube Music
图 1 · 摘自论文原文
  • 在推荐系统各层尝试六种干预策略,聚焦服务、训练、架构和探索四层。
  • 基于SNGP的不确定性探索使新歌推荐提升最明显,但牺牲部分用户参与度。
  • 服务层干预易被学习循环抵消,架构去偏虽增多样性却有集成成本。

持续训练的音乐推荐模型会陷入反馈循环,导致已听内容主导推荐结果,抑制新发布作品(时间新鲜度)和未听曲目(新颖性)的曝光。本文在YouTube Music首页开展离策略在线A/B测试,评估六种干预措施及其组合在四个概念层级(服务、训练、架构、探索)的效果。所有干预均作用于排序模型或其服务层,上游候选生成等环节保持不变。结果显示:服务层干预在持续训练系统中被学习循环中和;架构去偏可缓解流行度主导问题并提升多样性,但无法真正促进发现,且存在隐性集成成本;基于谱归一化神经高斯过程(SNGP)的不确定性探索带来最大新歌推荐提升,但伴随明显的参与度或多样性权衡。最后提出各层级干预的适用建议及潜在代价。

原文摘要 · Abstract (English)

Continuously trained ranking models in music recommenders fall into feedback loops where previously consumed items dominate recommendations. This suppresses two distinct content classes: new releases (temporal freshness) and unlistened catalog items (novelty). Industry practitioners have a wide menu of interventions available, ranging from serving-time heuristics, training-data reweighting, architectural debiasing, to uncertainty-driven exploration, each of which are well understood in academic settings. But live systems offer challenges with continuously ingested content, interconnected components, and practical limitations that counteract the findings from academic research. We report results from off-policy online A/B tests for six interventions and a combination experiment across four conceptual layers (serving, training, architecture, exploration) on the YouTube Music homepage. All interventions modify the ranking model or the serving layer that consumes its scores; candidate generation and other upstream components are held fixed. We discuss key takeaways from our results: first, serving-time interventions on continuously trained systems are neutralized by the learning loop. Second, architectural debiasing reduces popularity dominance and improves diversity but does not create discovery, while carrying hidden integration costs. Finally, uncertainty-driven exploration interventions with a Spectral-normalized Neural Gaussian Process (SNGP) head produce the largest new-release lift, though they come with a measurable engagement or diversity tradeoff. We close with recommendations on which layer to intervene at, and the hidden costs of each choice.

推荐系统去偏新歌发现A/B测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。