用大模型理解用户意图,让推荐更精准多样。
RecGPT Technical Report
- 以用户意图为核心,整合大模型于兴趣挖掘、召回与解释生成
- 在淘宝上线后提升内容多样性与用户满意度,各方获益
- 通过人机协作训练框架,实现大模型在推荐场景的高效对齐
推荐系统是人工智能最具影响力的应用之一,作为连接用户、商家与平台的关键基础设施。然而,当前工业系统仍严重依赖历史共现模式与日志拟合目标,即仅优化过去用户行为,未显式建模用户意图。这种日志拟合方法常导致对狭窄历史偏好的过拟合,难以捕捉用户动态演变和潜在兴趣,进而强化信息茧房与长尾现象,损害用户体验并威胁推荐生态可持续性。为此,我们重新思考推荐系统整体设计范式,提出RecGPT——下一代以用户意图为中心的推荐框架。通过将大语言模型(LLMs)融入用户兴趣挖掘、物品召回与解释生成等关键阶段,RecGPT将日志拟合推荐转化为意图驱动过程。为在大规模上有效对齐通用大模型与领域特定任务,RecGPT采用多阶段训练范式,融合推理增强预对齐与自训练演化,并由人机协同判别系统引导。目前,RecGPT已在淘宝App全面部署。线上实验表明,其在各利益相关方均实现持续性能提升:用户获得更高内容多样性与满意度,商家与平台获得更多曝光与转化。跨多方的综合改善验证了大模型驱动、意图中心的设计可构建更可持续、互利共赢的推荐生态系统。
原文摘要 · Abstract (English)
Recommender systems are among the most impactful applications of artificial intelligence, serving as critical infrastructure connecting users, merchants, and platforms. However, most current industrial systems remain heavily reliant on historical co-occurrence patterns and log-fitting objectives, i.e., optimizing for past user interactions without explicitly modeling user intent. This log-fitting approach often leads to overfitting to narrow historical preferences, failing to capture users' evolving and latent interests. As a result, it reinforces filter bubbles and long-tail phenomena, ultimately harming user experience and threatening the sustainability of the whole recommendation ecosystem. To address these challenges, we rethink the overall design paradigm of recommender systems and propose RecGPT, a next-generation framework that places user intent at the center of the recommendation pipeline. By integrating large language models (LLMs) into key stages of user interest mining, item retrieval, and explanation generation, RecGPT transforms log-fitting recommendation into an intent-centric process. To effectively align general-purpose LLMs to the above domain-specific recommendation tasks at scale, RecGPT incorporates a multi-stage training paradigm, which integrates reasoning-enhanced pre-alignment and self-training evolution, guided by a Human-LLM cooperative judge system. Currently, RecGPT has been fully deployed on the Taobao App. Online experiments demonstrate that RecGPT achieves consistent performance gains across stakeholders: users benefit from increased content diversity and satisfaction, merchants and the platform gain greater exposure and conversions. These comprehensive improvement results across all stakeholders validates that LLM-driven, intent-centric design can foster a more sustainable and mutually beneficial recommendation ecosystem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。