arXiv:2512.24325cs.IRcs.LG2025-12

用多智能体强化学习动态分配推荐系统算力,提升广告收入16.67%。

MaRCA: Multi-Agent Reinforcement Learning for Dynamic Computation Allocation in Large-Scale Recommender Systems

  • 将推荐系统各阶段视为协作智能体,采用集中训练分散执行策略优化算力分配。
  • 在日均百亿级请求下,仅用现有资源实现16.67%的广告收入增长。
  • 适用于大规模推荐系统算力优化,尤其适合高流量电商广告场景。

现代推荐系统因模型复杂度和流量规模上升面临重大计算挑战,高效算力分配对最大化商业收益至关重要。现有方法通常简化多阶段算力分配,忽略阶段间依赖,限制全局最优性。本文提出MaRCA,一种面向大规模推荐系统的端到端算力分配多智能体强化学习框架。MaRCA将推荐系统各阶段建模为协同智能体,采用集中训练、分散执行(CTDE)机制,在算力约束下优化收益。引入AutoBucket TestBench进行精准计算成本估算,并设计基于模型预测控制(MPC)的收益-成本平衡器,主动预测流量负载并动态调整收益与成本权衡。自2024年11月在某全球领先电商平台广告管道中端到端部署以来,MaRCA持续处理每日百亿级广告请求,仅使用现有算力即实现16.67%的收入提升。

原文摘要 · Abstract (English)

Modern recommender systems face significant computational challenges due to growing model complexity and traffic scale, making efficient computation allocation critical for maximizing business revenue. Existing approaches typically simplify multi-stage computation resource allocation, neglecting inter-stage dependencies, thus limiting global optimality. In this paper, we propose MaRCA, a multi-agent reinforcement learning framework for end-to-end computation resource allocation in large-scale recommender systems. MaRCA models the stages of a recommender system as cooperative agents, using Centralized Training with Decentralized Execution (CTDE) to optimize revenue under computation resource constraints. We introduce an AutoBucket TestBench for accurate computation cost estimation, and a Model Predictive Control (MPC)-based Revenue-Cost Balancer to proactively forecast traffic loads and adjust the revenue-cost trade-off accordingly. Since its end-to-end deployment in the advertising pipeline of a leading global e-commerce platform in November 2024, MaRCA has consistently handled hundreds of billions of ad requests per day and has delivered a 16.67% revenue uplift using existing computation resources.

推荐系统强化学习算力分配多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。