arXiv:2512.14503cs.IRcs.CL2025-12被引 8

用多智能体系统提升推荐模型效率与解释力,实测点击率提升3%。

RecGPT-V2 Technical Report

  • 构建分层多智能体协同推理,减少重复计算并覆盖多样意图。
  • 压缩用户行为表征,降低60%显存消耗,独占召回率升至10.99%。
  • 引入动态提示与智能体评鉴机制,解释更贴近人类偏好。

大型语言模型在推荐系统中展现出从隐式行为匹配转向显式意图推理的潜力。尽管RecGPT-V1已初步实现这一范式,但仍存在四大局限:(1)多路径推理导致计算低效与认知冗余;(2)固定模板生成解释多样性不足;(3)监督学习下泛化能力有限;(4)以结果为导向的评估无法匹配人类标准。为此,我们提出RecGPT-V2,包含四项核心创新:第一,采用分层多智能体系统重构意图推理,通过协同合作消除认知重复,并结合混合表征推理压缩用户行为上下文,使GPU消耗降低60%,独占召回率从9.39%提升至10.99%;第二,设计元提示框架动态生成情境自适应提示,解释多样性提升7.3%;第三,采用约束强化学习缓解多奖励冲突,在标签预测上提升24.1%,解释接受度提升13.0%;第四,提出智能体作为裁判框架,将评估分解为多步推理过程,显著提升与人类偏好的对齐。在淘宝线上A/B测试中,点击率(CTR)提升2.98%,独立访问量(IPV)提升3.71%,转化率(TV)提升2.19%,新用户留存率(NER)提升11.46%。RecGPT-V2验证了大规模部署基于LLM的意图推理在技术和商业上的可行性,弥合认知探索与工业应用之间的鸿沟。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated remarkable potential in transforming recommender systems from implicit behavioral pattern matching to explicit intent reasoning. While RecGPT-V1 successfully pioneered this paradigm by integrating LLM-based reasoning into user interest mining and item tag prediction, it suffers from four fundamental limitations: (1) computational inefficiency and cognitive redundancy across multiple reasoning routes; (2) insufficient explanation diversity in fixed-template generation; (3) limited generalization under supervised learning paradigms; and (4) simplistic outcome-focused evaluation that fails to match human standards. To address these challenges, we present RecGPT-V2 with four key innovations. First, a Hierarchical Multi-Agent System restructures intent reasoning through coordinated collaboration, eliminating cognitive duplication while enabling diverse intent coverage. Combined with Hybrid Representation Inference that compresses user-behavior contexts, our framework reduces GPU consumption by 60% and improves exclusive recall from 9.39% to 10.99%. Second, a Meta-Prompting framework dynamically generates contextually adaptive prompts, improving explanation diversity by +7.3%. Third, constrained reinforcement learning mitigates multi-reward conflicts, achieving +24.1% improvement in tag prediction and +13.0% in explanation acceptance. Fourth, an Agent-as-a-Judge framework decomposes assessment into multi-step reasoning, improving human preference alignment. Online A/B tests on Taobao demonstrate significant improvements: +2.98% CTR, +3.71% IPV, +2.19% TV, and +11.46% NER. RecGPT-V2 establishes both the technical feasibility and commercial viability of deploying LLM-powered intent reasoning at scale, bridging the gap between cognitive exploration and industrial utility.

推荐系统多智能体意图推理大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。