让评估贯穿LLM代理全生命周期,实现持续迭代与安全演化
Evaluation-Driven Development and Operations of LLM Agents: A Process Model and Reference Architecture
- 构建评估驱动的开发运维流程,将评估嵌入闭环反馈
- 融合离线与在线评估,支持系统运行中动态调整
- 适合需长期演进的生产级LLM代理团队参考
大型语言模型(LLMs)催生了能够追求模糊目标并在部署后自适应的智能体系统。由于其行为具有开放性、概率性和随时间演化的系统级交互特征,传统基于固定基准和静态测试集的评估方法难以捕捉涌现行为或支撑全生命周期的持续适应。为此,我们通过多视角文献综述(MLR)整合学术界与工业界的评估实践,提炼出两个实证衍生成果:一个流程模型与一个参考架构。二者共同构成评估驱动的开发与运维(EDDOps)方法,将评估作为持续性的核心控制机制而非最终检查点。该方法在闭环反馈中统一了开发期(offline)与运行期(online)评估,使评估结果同时驱动运行时自适应与受控式重构,从而支持更安全、可追溯的演化,确保LLM智能体始终对齐不断变化的目标、用户需求与治理约束。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have enabled the emergence of LLM agents, systems capable of pursuing under-specified goals and adapting after deployment. Evaluating such agents is challenging because their behavior is open ended, probabilistic, and shaped by system-level interactions over time. Traditional evaluation methods, built around fixed benchmarks and static test suites, fail to capture emergent behaviors or support continuous adaptation across the lifecycle. To ground a more systematic approach, we conduct a multivocal literature review (MLR) synthesizing academic and industrial evaluation practices. The findings directly inform two empirically derived artifacts: a process model and a reference architecture that embed evaluation as a continuous, governing function rather than a terminal checkpoint. Together they constitute the evaluation-driven development and operations (EDDOps) approach, which unifies offline (development-time) and online (runtime) evaluation within a closed feedback loop. By making evaluation evidence drive both runtime adaptation and governed redevelopment, EDDOps supports safer, more traceable evolution of LLM agents aligned with changing objectives, user needs, and governance constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。