让评测基准自动进化,生成更难且可验证的智能体任务
Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
- 通过智能体自由探索,动态生成更高难度的新任务
- 在GAIA和AIME-2024上均提升任务复杂度与结果可靠性
- 适合关注智能体评估体系可持续发展的研究者
大型语言模型和智能体系统的发展使智能体具备前所未有的能力,但现有评测基准正迅速被新智能体突破上限,难以满足评估需求。为此,我们提出轨迹式可复现验证的智能体基准复杂度演化框架(TRACE)。该框架从原始任务出发,引导智能体自由探索并演化出更难的新任务,同时记录可验证的执行轨迹。流程包括:(1) 演化提案挖掘,通过初步探索和发散思维生成任务演化方案;(2) 问题构建与自由探索,将提案转化为可行问题并让智能体自由探索,记录执行轨迹;(3) 多层级验证,确保演化后任务具有可验证、可复现的轨迹。在GAIA基准上的实验表明,TRACE持续提升任务复杂度,并通过可验证轨迹增强正确性可靠性。此外,框架成功适配并提升了以AIME-2024为代表的推理数据集。本工作标志着从静态人工构建基准向动态自演化评估系统的范式转变,为智能体发展提供可持续且具挑战性的评测跑道。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) and agent system designs have empowered agents with unprecedented levels of capability. However, existing agent benchmarks are showing a trend of rapid ceiling-hitting by newly developed agents, making it difficult to meet the demands for evaluating agent abilities. To address this problem, we propose the Trajectory-based Validated-by-Reproducing Agent-benchmark Complexity Evolution (TRACE) framework. This framework takes an original task from an existing benchmark and encourages agents to freely explore and evolve it into a new task with higher difficulty while recording validatable agent trajectories. The framework proceeds in three stages: (1) evolutionary proposal mining, which provides task evolution proposals through preliminary exploration and divergent thinking; (2) problem formation and free exploration, where proposals are conceptualized into feasible problem candidates and the agents then explore them freely while recording their execution trajectories; and (3) multi-level validation, which ensures that the evolved tasks are accompanied by validatable and reproducible trajectories. Experiments on the GAIA benchmark demonstrate that the TRACE framework consistently enhances task complexity while improving the reliability of correctness through validatable execution trajectories. In addition, our framework can successfully adapt to and improve reasoning datasets represented by AIME-2024. This work marks a paradigm shift from static, manually curated benchmarks to dynamic, self-evolving evaluation systems, providing a sustainable and challenging runway for agent development
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。