arXiv:2602.18640cs.AI2026-02被引 2

让大模型自动优化推荐系统,用智能体推理提升效果与稳定性。

Decoding ML Decision: An Agentic Reasoning Framework for Large-Scale Ranking System

  • 用可编程环境中的智能体自主探索最优排序策略。
  • 在多个产品场景中找到近似帕累托最优的高效策略。
  • 适合需要快速迭代、保证上线稳定的推荐系统团队。

现代大规模排序系统面临多重目标冲突、运行约束和不断变化的产品需求。该领域的进展日益受制于工程上下文约束:将模糊的产品意图转化为合理、可执行、可验证的假设,远比建模技术本身更难。我们提出GEARS(生成式智能体排序系统引擎),将排序优化重构为可编程实验环境中的自主发现过程。不同于静态的模型选择,GEARS通过专用智能体技能将排序专家知识封装为可复用的推理能力,使运营者可通过高层意图感知实现个性化调控。此外,为保障生产可靠性,框架内置验证钩子,确保统计稳健性并过滤过拟合短期信号的脆弱策略。在多种产品场景的实验验证表明,GEARS能持续识别出性能更优、接近帕累托效率的策略,同时融合算法信号与深层排序上下文,并保持严格的部署稳定性。

原文摘要 · Abstract (English)

Modern large-scale ranking systems operate within a sophisticated landscape of competing objectives, operational constraints, and evolving product requirements. Progress in this domain is increasingly bottlenecked by the engineering context constraint: the arduous process of translating ambiguous product intent into reasonable, executable, verifiable hypotheses, rather than by modeling techniques alone. We present GEARS (Generative Engine for Agentic Ranking Systems), a framework that reframes ranking optimization as an autonomous discovery process within a programmable experimentation environment. Rather than treating optimization as static model selection, GEARS leverages Specialized Agent Skills to encapsulate ranking expert knowledge into reusable reasoning capabilities, enabling operators to steer systems via high-level intent vibe personalization. Furthermore, to ensure production reliability, the framework incorporates validation hooks to enforce statistical robustness and filter out brittle policies that overfit short-term signals. Experimental validation across diverse product surfaces demonstrates that GEARS consistently identifies superior, near-Pareto-efficient policies by synergizing algorithmic signals with deep ranking context while maintaining rigorous deployment stability.

排序系统智能体推荐优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。