用多智能体自动迭代推荐系统,实现自我进化。
AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems

- 四阶段闭环:构思、编码、评估、进化,全自动化
- 自动生成实验并上线,实验速度远超人工
- 适合需要持续优化推荐系统的工业级团队
推荐算法迭代正从依赖工程师的手工流程转向工业化研究循环,但这一转变仍受制于结构性执行瓶颈:从想法到上线仍需人工生成假设、修改生产代码、开展A/B实验并归因结果。创新因此仅随人力线性增长,而非基于证据、算力和积累的实验知识复利式提升。我们提出AgentX,一个已投入生产的多智能体系统,从根本上重构了这一生产函数。AgentX作为自演进开发引擎,可自主生成、实施、评估并从推荐实验中学习,其规模与速度远超人工工作流。系统在闭环中协调四个紧密耦合阶段:头脑风暴代理综合历史实验、系统架构、数据分析和外部研究,生成排序后的可执行提案;开发代理通过仓库驱动生成和多维可靠性验证,将提案转化为生产就绪代码;评估代理在安全护栏控制下进行线上灰度发布,通过有否决权的A/B判断将成功与失败转化为结构化知识资产;演化层(SGPO)则从执行轨迹中提炼语义梯度更新,持续优化各代理自身——使系统不仅自动化,更具备自我改进能力。
原文摘要 · Abstract (English)
Recommendation algorithm iteration is moving from an artisanal, engineer-bound process toward an industrialized research loop, but this transition remains blocked by a structural execution bottleneck: the idea-to-launch cycle still depends on human engineers to generate hypotheses, modify production code, launch A/B experiments, and attribute online results. Innovation therefore scales linearly with headcount rather than compounding with evidence, compute, and accumulated experimental knowledge. We present AgentX, a production-deployed multi-agent system that fundamentally restructures this production function. AgentX operates as a self-evolving development engine: it autonomously generates, implements, evaluates, and learns from recommendation experiments at a scale and pace that no manual workflow can sustain. The system orchestrates four tightly coupled stages in a closed loop. A Brainstorm Agent synthesizes evidence from historical experiments, system architecture, data analysis, and external research into ranked, executable proposals. A Developing Agent translates each proposal into production-ready code through repository-grounded generation and multi-dimensional reliability verification. An Evaluation Agent conducts safe online rollout with guardrail-vetoed A/B judgment, converting both successes and failures into structured knowledge assets. A Harness Evolution layer (SGPO) then distills execution trajectories into semantic-gradient updates that continuously sharpen the agents themselves -- making the system not merely automated, but self-improving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。