让大模型助手自动优化自身能力,更稳定高效。
DREvo: Distilling Recalibrated Historical Experience for Harness Self-Evolution

- 用动态校准历史经验,判断哪些过往数据仍可用。
- 在有限试错次数下,五项任务准确率最高,平均提升超14%。
- 适合需要持续进化的大模型智能体开发者使用。
大语言模型智能体的性能高度依赖其控制机制(harness),而构建高性能的 harness 需要大量专家投入。近年来研究转向 harness 自我演化,通过迭代提出、评估和改进 harness 来利用历史试错经验。然而,累积的历史经验常无法提供稳定搜索指引,导致性能在演化过程中波动剧烈,在有限演化预算下难以可靠发现高性能 harness。我们识别出现有方法的两大缺陷:(1) 缺乏对历史经验是否仍适用于当前 harness 的动态再评估;(2) 缺少将有效经验转化为可操作演化方向的显式机制。为此,我们提出 DREvo,融合功能级证据锚定、状态依赖证据校准与角色条件化搜索意图蒸馏,以判断哪些历史经验仍有效,并确定下一步演化方向。在有限演化预算下,DREvo 展现出更平滑的演化轨迹,在全部五个基准测试中达到最高准确率,于领域推理和代理任务上分别相较基线平均提升 16.2% 和 14.2%。
原文摘要 · Abstract (English)
Harness plays a critical role in large language model agent performance, and building a high-performing harness requires substantial expert effort. Therefore, recent research has increasingly explored harness self-evolution, which iteratively proposes, evaluates, and improves harnesses using historical trial experience. However, accumulated historical experience does not always translate into stable search guidance, and performance often fluctuates substantially across evolution iterations, making it difficult to reliably discover high-performing harnesses under a limited evolution budget. We identify two limitations in how existing harness self-evolution methods leverage historical experience: (1) Lack of dynamic reassessment of whether historical experience remains valid for the current harness, and (2) Lack of explicit mechanisms for translating valid historical experience into actionable search directions. To address these limitations, we propose a new harness self-evolution method, named DREvo, which integrates function-level evidence anchoring, state-dependent evidence recalibration, and role-conditioned search intent distillation to determine which historical evidence remains valid and where the harness should evolve next. Under limited evolution budgets, DREvo exhibits smoother evolution trajectories, achieves the highest accuracy on all five benchmarks, and delivers average gains of 16.2% and 14.2% over the evaluated baselines on domain reasoning and agentic tasks, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。