arXiv:2606.06114cs.AI2026-06被引 1

用人类监督模拟器防止自进化智能体能力退化和安全漂移

Towards Healthy Evolution: Exploring the Role and Mechanisms of Human-Agent Interaction in Self-Evolving Systems

论文配图:Towards Healthy Evolution: Exploring the Role and Mechanisms of Human-Agent Interaction in Self-Evolving Systems
图 1 · 摘自论文原文
  • 设计基于大模型的监督框架,模拟人类在自进化各阶段提供反馈
  • 有限监督可显著缓解安全问题,同时保持核心任务性能稳定
  • 输出验证阶段干预最有效,频繁监督收益递减,适合工程部署

自进化智能体通过持续自我对弈和自生成学习信号提升能力,但自主进化可能引发能力退化与安全漂移。尽管人类反馈在静态或训练后模型中已证明有效,其在自进化系统中的作用仍不明确。本文提出基于大语言模型的代理规范纠正框架(ANCHOR),模拟人类监督,在自进化不同阶段提供反馈。我们在编码、数学推理与安全三个任务上评估两个开源自进化代理系统。结果表明,即使有限监督也能显著缓解安全退化,同时维持核心进化目标的稳定性能。进一步分析显示,对输出验证阶段的干预最为有效,而提高监督频率带来的边际收益递减。这些发现为设计更稳定、可控且与人类对齐的自进化系统提供了实证支持与实践指导。

原文摘要 · Abstract (English)

Self-evolving agents improve through continual self-play and self-generated learning signals, but autonomous evolution can also cause capability degradation and safety drift. Although human feedback has proven effective for static and post-trained agents, its role in self-evolving systems remains underexplored. We introduce Agent Norm Correction through Human-like Oversight and Review (ANCHOR), an LLM-based framework that simulates human supervision and delivers feedback at various phases of self-evolution. With ANCHOR, we evaluate two representative open-source self-evolving agent systems across coding, mathematical reasoning, and safety. Our results show that even limited supervision substantially mitigates safety degradation while preserving stable performance on core evolutionary objectives. Further analysis shows that supervision over the output verification phase is the most effective for intervention, whereas increasing supervision frequency yields diminishing returns. These findings provide empirical evidence and practical guidance for designing more stable, controllable, and human-aligned self-evolving agent systems.

自进化人机交互安全控制大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。