分析安全提示自进化机制的可行性与极限,揭示其改进瓶颈。
Safe Harness Self-Evolution: A Theoretical Analysis of Feasibility and Limits
- 通过理论推导建立奖励提升条件,控制任务保留变化
- 证明生成合格修改的概率与当前表现无关,评估受限时增加候选不保证成功
- 揭示持续改进可能停滞,适合研究自进化安全机制的学者
提示自进化指代理在保持底层语言模型冻结的前提下,根据任务反馈动态修改提示、工具、代码或编排策略,且变更可跨任务持久化。本文对安全提示自进化提供系统性理论分析,关联修改生成、有限数据认证与选择、安全采纳及更新后行为。在固定用户-任务分布下,我们建立了保障整体期望奖励提升的同时控制保留任务变化的条件,刻画了生成合格修改的概率,并推导出安全选择与采纳的有限数据边界。分析表明,生成与认证施加不同约束:当前任务表现不决定生成合格修改的概率,且当评估受限时,增加候选数量未必提高更新成功的保证。因此,即使存在改进机会,也可能陷入停滞。进一步发现,识别真实改进的最坏情况评估成本在期望奖励趋近上界时发散。连续更新中,认证改进保证可在有限轮次内累积,但一次成功更新本身并不保证后续仍可继续改进。这些结果为诊断瓶颈和设计更安全的自进化机制提供了理论基础。
原文摘要 · Abstract (English)
Harness self-evolution is the process by which an agent modifies its prompts, tools, code, or orchestration in response to task feedback while keeping the underlying language model frozen, with changes persisting across subsequent tasks. We provide a systematic theoretical analysis of the feasibility and limits of safe harness self-evolution, connecting modification generation, finite-data certification and selection, safe adoption, and behavior after an update. Under a fixed user-task distribution, we establish conditions guaranteeing overall expected-reward improvement while controlling changes on retained tasks, characterize the probability of generating qualified modifications, and derive finite-data bounds for safe selection and adoption. Our analysis shows that generation and certification impose distinct constraints: current task performance does not determine the probability of generating qualified modifications, and generating more candidates need not improve the guarantee of a successful update when evaluation is limiting. Stagnation may therefore arise even when improvement opportunities remain. We further show that worst-case evaluation cost for recognizing genuine improvements diverges as expected reward approaches its upper bound. Across successive updates, certified improvement guarantees accumulate over a finite run, but a successful update does not by itself guarantee that further improvement remains possible. These results provide a basis for diagnosing bottlenecks and designing safer self-evolution mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。