用智能代理自动探索广告排序模型优化,提升迭代效率。
Agentic ML Exploration (A-MLE) for Ads Ranking
- 构建智能代理系统,自动化完成从假设到部署的完整机器学习迭代流程。
- 在多个大规模广告排序模型上验证,显著缩短优化周期,释放长期被忽视的性能潜力。
- 适合需要持续优化但人力稀缺的工业推荐系统团队使用。
现代工业级广告排序系统的主要瓶颈已不再是模型容量或训练算力,而是人工机器学习迭代的吞吐量——每项统计显著改进需经历研究、实现、训练、调试、评估和上线等多轮耗时数天至数周的人工参与。典型排序系统包含众多差异化模型,数据、架构与基础设施约束各异,每个模型每次迭代均需资深工程师数日到数周投入。导致已在某模型有效的技术难以快速、均衡地推广至其他模型,大量可恢复信号未被发掘。本文提出自主式机器学习探索系统 A-MLE,通过一个智能体协调五阶段流程:假设生成、探索策略制定、实验执行、结果分析与知识共享,结合领域特定技能与代理工作流,在沙箱环境中运行,并在每阶段边界设置人机协同检查点。我们在代表性大规模广告排序模型上部署 A-MLE,基于工具可用性、自主工作流执行与开放式探索三个层级评估其能力。进一步开展固定代理循环的跨大模型对比研究,揭示 Claude Sonnet、Gemini 与 GPT 系列在执行可靠性与探索激进性上的定性差异。讨论失败模式及影响可靠性的设计选择。结果表明,代理式探索是工业推荐系统中提升机器学习工程师效能的实际倍增器,尤其适用于极少获得专家关注的长尾模型。
原文摘要 · Abstract (English)
Modern industrial ads ranking stacks are increasingly bottlenecked not by model capacity or training compute, but by the throughput of human ML iteration - the cycles of research, implementation, training, debugging, evaluation, and launch required to surface a single statistically significant improvement. A typical ranking stack contains numerous differentiated models with heterogeneous data, architectures, and infrastructure constraints, and each cycle takes days to weeks of senior engineer attention per model. As a result, techniques that have proven effective on one model diffuse into others slowly and unevenly, leaving substantial recoverable signal unexplored. We present Agentic ML Exploration (A-MLE), an autonomous LLM-agent system that systematically explores ML techniques across a portfolio of ads ranking models. A-MLE decomposes ML iteration into five stages involving hypothesis generation, exploration strategy, experiment execution, result analysis and shared knowledge substrate which are orchestrated by a single agent that invokes domain-specific skills and agentic workflows against a sandboxed execution layer, with human-in-the-loop checkpoints at each stage boundary. We deploy A-MLE across a representative set of large-scale ads ranking models and evaluate it along a tiered capability framework (tool availability, autonomous workflow execution, and open-ended exploration). We further report a controlled cross-LLM study using a fixed agent loop, which surfaces qualitative differences in execution reliability and exploration aggressiveness across the Claude Sonnet, Gemini, and GPT families. We discuss failure modes and the design choices that govern reliability. Our findings suggest that agentic exploration is a practical force multiplier for ML engineers in industrial recommenders, especially for the long tail of models that rarely receive expert attention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。