用智能演化方法在多个价格点上击败所有竞品模型。
Competing at Every Price Point with Agentic Evolution over a Menu of LLMs
- 通过进化算法在9个不同价格的LLM中自动组合生成最优代理程序。
- 在两个任务上实现全价格区间性能领先,包括最高分和最低成本点。
- 仅需100个训练样本,适合想低成本打造顶尖智能系统的团队。
一家公司希望在特定智能任务上,以所有竞争对手的价格点提供更优的准确率。若能实现对对手的帕累托支配,则无理性客户会转向其他产品。本文提出一种基于智能体演化的路径,在仅需最多100个训练示例的情况下,从一个包含9个不同定价的LLM接口的菜单中,通过一个简单种子代理与用户设定的成本目标(通常为现有竞品价格),由名为RoboPhD的进化型元智能体逐步优化出完整的代理程序。该方法在两个语义差异较大的任务上持续突破:DS-1000(执行验证的代码生成)与PaperFindingBench(LLM评估的科学文献检索)。其官方评分提交结果占据两个榜单上的所有帕累托前沿位置(除一处外),并实现了对最高分与最低成本竞争点的全面超越。
原文摘要 · Abstract (English)
Consider a firm that surveys its competition for a particular agentic task and seeks to offer superior accuracy at every competitor price point. A firm that Pareto-dominated its competitors would leave no rational customer a reason to buy elsewhere. This paper shows a path to this kind of capability via agentic evolution over a menu of LLMs, from training pools of at most 100 examples. Given a priced menu of nine LLM endpoints; brief documentation of the task, objective, and API; a simple seed agent; and an operator-chosen per-problem cost target - usually set at an incumbent's own price - RoboPhD, an evolutionary meta-agent, evolves complete agent programs that attack the public frontiers of two semantically dissimilar tasks point by point: DS-1000 (execution-checked code generation) and PaperFindingBench (LLM-judged scientific document retrieval). Our officially scored submissions hold every Pareto-frontier slot but one on the two tasks' leaderboards, including Pareto domination of both the top-scoring and the lowest-cost competing points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。