让AI Agent自我进化更安全可审计,把改进变成可验证的技能积累。
Audited Skill-Graph Self-Improvement for Agentic LLMs via Verifiable Rewards, Experience Synthesis, and Continual Memory
- 将每次改进转化为有接口的可复用技能,通过验证器审核后才纳入系统。
- 用可回放的证据分解奖励,确保每步优化都能被独立审计。
- 适合关注AI安全、可解释性及持续演化的研究者和工程师。
强化学习正被用于将大语言模型转化为具备长期规划能力的智能体,能调用工具、管理记忆并在部分可观测环境下行动。尽管近期工作在工具学习、可验证奖励与持续训练方面取得进展,但自进化智能体仍面临未解决的安全与治理挑战:优化压力可能诱发奖励欺骗,行为漂移难以审计或复现,且改进常隐含于不透明的参数更新中而非可重用的成果。本文提出审计型技能图自提升框架(ASG-SI),将自提升视为智能体逐步编译为不断增长、可审计的技能图过程。每个候选改进从成功轨迹中提取,标准化为具明确接口的技能,并在通过验证器支持的回放测试与合约检查后才被采纳。奖励被分解为可重构的组件,基于可回放证据生成,实现对晋升决策与学习信号的独立审计。ASG-SI进一步整合经验合成以实现规模化压力测试,并通过持续记忆控制维持有限上下文下的长时性能。本文提出完整系统架构、威胁模型与安全分析,并提供可运行的参考实现,展示验证器支持的奖励构建、技能编译、审计日志记录及在持续任务流中的可测量改进。ASG-SI将智能体自提升重新定义为可验证、可重用能力的积累,为可复现评估与运营治理自进化AI智能体提供了可行路径。
原文摘要 · Abstract (English)
Reinforcement learning is increasingly used to transform large language models into agentic systems that act over long horizons, invoke tools, and manage memory under partial observability. While recent work has demonstrated performance gains through tool learning, verifiable rewards, and continual training, deployed self-improving agents raise unresolved security and governance challenges: optimization pressure can incentivize reward hacking, behavioral drift is difficult to audit or reproduce, and improvements are often entangled in opaque parameter updates rather than reusable, verifiable artifacts. This paper proposes Audited Skill-Graph Self-Improvement (ASG-SI), a framework that treats self-improvement as iterative compilation of an agent into a growing, auditable skill graph. Each candidate improvement is extracted from successful trajectories, normalized into a skill with an explicit interface, and promoted only after passing verifier-backed replay and contract checks. Rewards are decomposed into reconstructible components derived from replayable evidence, enabling independent audit of promotion decisions and learning signals. ASG-SI further integrates experience synthesis for scalable stress testing and continual memory control to preserve long-horizon performance under bounded context. We present a complete system architecture, threat model, and security analysis, and provide a fully runnable reference implementation that demonstrates verifier-backed reward construction, skill compilation, audit logging, and measurable improvement under continual task streams. ASG-SI reframes agentic self-improvement as accumulation of verifiable, reusable capabilities, offering a practical path toward reproducible evaluation and operational governance of self-improving AI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。