让大模型推荐更懂用户,兼顾语义理解与点击偏好。
Taiji: Pareto Optimal Policy Optimization with Semantics-IDs Trade-off for Industrial LLM-Enhanced Recommendation

- 用逆向推理和拒绝采样生成高质量推理数据
- 提出帕累托最优策略优化,动态平衡语义与偏好奖励
- 已在快手广告平台服务超4亿日活用户
通过大语言模型(LLM)扩展推荐系统已成为工业界主流趋势。然而,如何在后训练阶段(如SFT和RL)将LLM的语义空间与推荐系统的ID空间对齐仍具挑战。现有LLM4Rec范式受两大问题制约:(1) 在开放域推荐中难以衡量与提升思维链(CoT)质量;(2) 强化学习对齐过程中忽视了LLM语义奖励与推荐偏好奖励之间的权衡。受此启发,我们提出Taiji,一种面向工业推荐系统的新型大模型增强框架。为突破SFT瓶颈,采用逆向工程推理与开放式拒绝采样生成高质量、领域相关的CoT数据;为解决RL对齐难题,提出帕累托最优策略优化(POPO),自适应调整跨域奖励权重。理论上实现了大模型语义知识与协同ID特征所代表的在线用户偏好之间的最优权衡。大量离线评估与线上A/B测试验证了其有效性。自2026年5月起部署于快手广告平台,目前每日服务超4亿用户,带来显著商业收益,展现出在万级规模环境下的强可扩展性。
原文摘要 · Abstract (English)
Scaling recommender systems via large language models (LLMs) has become a prominent trend in the industry. However, aligning the LLM's semantic space with the recommender's ID space via post-training (e.g., SFT and RL) remains challenging. Existing LLM4Rec paradigms are bottlenecked by two main issues: (1) the difficulty of measuring and improving chain-of-thought (CoT) quality in open-domain recommendation during SFT, and (2) the neglect of the trade-off between LLM semantic rewards and recommendation preference rewards during RL alignment. Inspired by these challenges, we present Taiji, a novel LLM-as-Enhancer framework designed for industrial recommender systems. To overcome the SFT bottleneck, we utilize reverse-engineered reasoning and open-ended rejection sampling to generate high-quality, domain-specific CoT data. To resolve the RL alignment issue, we propose Pareto Optimal Policy Optimization (POPO), which adaptively adjusts cross-domain reward weights. Theoretically, it achieves an optimal trade-off between the semantic world knowledge of LLMs and the collaborative ID features representing online user preferences. Extensive offline evaluations and online A/B tests validate the effectiveness of Taiji. Deployed on Kuaishou's advertising platform since May 2026, Taiji currently serves over 400 million users daily, yielding significant commercial revenue and demonstrating its robust scalability in web-scale environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。