自评导致代理循环误判停滞为进展,需外部验证才能避免虚假进步。
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops
- 用隔离环境的全局状态真值作为评估基准,检验代理自评可靠性。
- 54轮中44%实际退步却被自评认为进步,最佳状态因自评衰减19%。
- 对开放目标,仅强化自评无效,必须引入外部真实世界验证机制。
长时间运行的自主代理在无外界干预下自主规划、执行并自我评判成果。当代理自行评估时,自评偏差产生:看似合理的改变被当作进展,而真实结果却停滞或倒退。我们称此为‘进展幻象’,并通过受控实验表明,问题根源在于评估者是否具备真实依据。构建测试平台,固定代理与工具界面,仅改变评估者的信息通道类型。采用容器与网络隔离的全局状态真值(world-state oracle),每轮验证不可伪造。在54轮循环中,前沿代理声称每轮都有改进,但56%的实际变化量为零或负值。自评结果完全失效,自评门禁退化为全接受,使最优状态下降19%。即使最强的内部评估者(读取完整文本、变更差异和历史判断)仍接受44%的真实退步,并拒绝38%的真实进步;预注册的对抗性假设‘强评估可弥合差距’被证伪。在成功标准可从输出物本身验证的边界任务中,同一评估者幻象消失,差距降至注册阈值内,证明差距取决于成功信号的来源位置。仅返回接受判定的简化版本,其真实输出与完整反馈相近(110.0 vs 113.0),说明收益来自评估门的接地性而非反馈内容。对于成功信号存在于文本之外的开放目标,单纯提升评估能力不足,必须采用具备真实世界访问能力的外部评估,这是结构上的必要条件。
原文摘要 · Abstract (English)
Long-running autonomous agents plan, act, and judge their own completion without human intervention. When an agent grades its own work, self-evaluation bias takes hold: plausible changes are accepted as progress while real-world outcomes stagnate or regress. We name this failure mode the progress mirage and show, with controlled measurement, that it is a question of what the evaluator is grounded in. We built a testbed that holds the agent and its tool surface fixed and manipulates only the information-channel type of the evaluator that gates the loop. A world-state oracle, unfakeable in principle, is enforced by container and network isolation and verified at every run. Across 54 cycles a frontier agent claimed improvement every time, yet 56 percent had a measured delta of zero or below. Self-report was thus uninformative, and the self-verdict gate degenerated into accept-all, eroding the best deployed state it had reached by 19 percent. Even the strongest in-band judge, reading the full artifact text, the change diff, and its own verdict history, accepted cycles of which 44 percent were real-world regressions and rejected 38 percent of real improvements; the preregistered adversarial hypothesis that a strong judge closes the gap was rejected. On a boundary task whose success specification is verifiable from the artifact itself, the same judge's mirage vanished to zero and the gap collapsed within the registered threshold, showing that the gap depends on where the success signal resides. A sign-only variant returning only the acceptance verdict kept real-world output similar to full feedback (110.0 versus 113.0), locating the benefit in the gate's grounding rather than in feedback content. For open-ended objectives whose success signal lives outside the transcript, scaling up the judge is not enough; out-of-band evaluation with real-world access is a structural requirement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。