arXiv:2607.02440cs.AIcs.CL2026-07

测试智能体在有限反馈下自我优化策略的能力,发现有效进化需匹配任务机制。

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

论文配图:EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments
图 1 · 摘自论文原文
  • 构建可控环境,让智能体在固定预算内迭代修改策略代码
  • GPT-5.5在16个环境中均表现顶尖,综合排名最高
  • 可追踪每步优化细节,揭示策略演化的实际路径

自主智能体越来越需要通过反馈改进可执行策略,但现有评估常将此过程简化为最终得分或与开放式的软件工程进展混同。我们提出自主策略演化(Autonomous Policy Evolution)这一受控评估框架,其中一个代理模型在固定交互预算下反复编辑可执行策略系统。我们在此框架下构建了EvoPolicyGym基准,基于紧凑的交互式强化学习环境,评估智能体如何迭代优化已探索的策略。在EvoPolicyGym套件中,GPT-5.5取得最强综合排名,在全部16个环境中均位列前二。除排行榜成绩外,EvoPolicyGym还提供轨迹级诊断,可区分智能体如何分配预算、将反馈转化为参数调优。分析显示,强的自主策略演化不仅依赖单任务成功,更依赖发现适配任务的机制,并在受限反馈下持续优化策略。

原文摘要 · Abstract (English)

Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it with open-ended software-engineering progress. We introduce Autonomous Policy Evolution, a controlled evaluation setting in which a harness-model agent repeatedly edits an executable policy system under a fixed interaction budget. We instantiate this setting in EvoPolicyGym, a benchmark built from compact interactive RL environments that evaluates how agents iteratively improve explored policies. On the EvoPolicyGym suite, GPT-5.5 achieves the strongest aggregate rank score and top-two performance on all 16 environments. Beyond leaderboard results, EvoPolicyGym also provides trajectory-level diagnostics that distinguish how agents allocate budget, convert feedback into parametric tuning. These analyses show that strong autonomous policy evolution depends not only on isolated task wins, but on discovering task-appropriate mechanisms and refining policies under bounded feedback.

自主进化策略优化评估基准RL

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。