arXiv:2602.15005cs.CLcs.IR2026-02

用强化学习生成用户兴趣查询,提升跨域新闻推荐效果

Learning User Interests via Reasoning and Distillation for Cross-Domain News Recommendation

  • 用大模型从多元行为中生成兴趣查询,优化推荐信号
  • 增加计算资源可线性提升推荐效果,具规模效应
  • 小模型蒸馏后仍保持高精度,适合大规模部署

新闻推荐在在线新闻平台中至关重要,帮助用户发现相关内容。跨域新闻推荐需从异构信号中推断用户的深层信息需求,超越表面行为以捕捉可复用的兴趣。本文提出一种强化学习框架,训练大语言模型从跨域用户信号生成高质量的兴趣驱动新闻搜索查询。将查询生成建模为策略优化问题,采用带有多种奖励信号的GRPO算法。系统研究了推理时采样与模型容量两个计算维度,实证观察到计算增加带来持续提升,表现出类似缩放规律。最后通过在线蒸馏,将大型高算力教师模型的策略迁移到轻量级学生模型,适用于大规模部署。大量离线实验、消融研究及生产环境中的大规模在线A/B测试表明,在兴趣建模质量和下游推荐性能上均取得稳定提升。

原文摘要 · Abstract (English)

News recommendation plays a critical role in online news platforms by helping users discover relevant content. Cross-domain news recommendation further requires inferring user's underlying information needs from heterogeneous signals that often extend beyond direct news consumption. A key challenge lies in moving beyond surface-level behaviors to capture deeper, reusable user interests while maintaining scalability in large-scale production systems. In this paper, we present a reinforcement learning framework that trains large language models to generate high-quality lists of interest-driven news search queries from cross-domain user signals. We formulate query-list generation as a policy optimization problem and employ GRPO with multiple reward signals. We systematically study two compute dimensions: inference-time sampling and model capacity, and empirically observe consistent improvements with increased compute that exhibit scaling-like behavior. Finally, we perform on-policy distillation to transfer the learned policy from a large, compute-intensive teacher to a compact student model suitable for scalable deployment. Extensive offline experiments, ablation studies and large-scale online A/B tests in a production news recommendation system demonstrate consistent gains in both interest modeling quality and downstream recommendation performance.

新闻推荐兴趣建模强化学习模型蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。