arXiv:2607.17719cs.AIcs.MA2026-07

用智能代理自动优化电商推荐排序策略,提升用户转化与体验。

SR-Agent: An Experience-Driven Agentic Framework for Post-Ranking Strategy Refinement in E-Commerce Recommendation

论文配图:SR-Agent: An Experience-Driven Agentic Framework for Post-Ranking Strategy Refinement in E-Commerce Recommendation
图 1 · 摘自论文原文
  • 构建三阶段智能体框架,自动发现并诊断推荐问题。
  • 线上测试中订单量提升0.71%,浏览深度增加0.34%。
  • 适合追求自动化优化推荐策略的工业级平台使用。

用户体验是工业级电商推荐系统的核心目标。后排序策略通过控制排序列表中的多样性、相似性与曝光分布,因其部署简便且服务成本低而广泛应用。然而,随着线上环境持续变化,静态配置的策略逐渐过时,导致体验下降。传统优化依赖人工排查与更新,效率低、成本高且难以复用。尽管现有基于大模型的智能体(如RecUserSim、SimUSER、Self-EvolveRec)提供了新方向,但尚未实现闭环的自动化策略自演化。为此,我们提出SR-Agent,据知是首个用于工业推荐系统后排序策略自动优化的智能体框架。该框架整合三个模块:(i) UserSim智能体,运用评估技能识别用户感知不佳的案例;(ii) Analysis智能体,将重复出现的问题归纳为结构化、可复用的诊断;(iii) 受限的策略优化引擎,将诊断映射为类型化、有界的操作,并通过四阶段奖励机制与可逆回滚保障安全。在快手电商平台上部署后,SR-Agent持续运行优化循环。一个月的在线A/B测试显示,订单量提升0.71%,浏览深度增加0.34%,点击类目多样性提高0.48%,同时显著缩短优化周期,降低运维成本。

原文摘要 · Abstract (English)

User experience is a first-class objective in industrial e-commerce recommender systems (RS). Post-ranking strategies, which govern diversity, similarity, and exposure over a ranked list, are widely deployed in industrial RS for their simplicity and low serving cost. However, as the online recommendation environment evolves continuously, these statically configured strategies gradually become stale, thereby degrading the user experience. Refining them typically relies on manual inspection, diagnosis, and updates, making it slow, costly, and difficult to scale or reuse. Although recent LLM-based agents (e.g., RecUserSim, SimUSER, and Self-EvolveRec) offer promising directions, none of them close the full loop of automated, self-evolving strategy refinement. To bridge this gap, we introduce SR-Agent, which, to the best of our knowledge, is the first agentic framework deployed to refine post-ranking strategies in industrial RS. SR-Agent unifies three components: (i) a UserSim agent that applies inspection skills to surface user-perceived bad cases; (ii) an Analysis agent that consolidates recurring bad cases into structured, reusable diagnoses; and (iii) a constrained Strategy Refinement Harness that maps diagnoses to typed and bounded actions, gated by a four-stage reward pipeline with reversible rollback. Deployed on the Kuaishou e-commerce platform, SR-Agent continuously runs this refinement loop and, in a one-month online A/B test, increases order volume by 0.71%, browsing depth by 0.34%, and clicked-category diversity by 0.48%, while markedly shortening the refinement cycle and lowering operational cost.

推荐系统智能体策略优化电商

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。