让大模型代理的失败变得可理解,降低人工排查成本。
VeriLA: A Human-Centered Evaluation Framework for Interpretable Verification of LLM Agent Failures
- 用人类设计的标准定义代理预期行为
- 训练人类对齐的验证模块自动评估执行结果
- 适合需高可靠性的复杂系统开发者使用
AI从业者越来越多地在复杂AI系统中使用大语言模型(LLM)代理解决推理任务,但这些代理的执行常因不符合人类标准而失败,导致系统整体性能下降。由于代理推理过程不透明、与人类预期错位、依赖关系复杂及人工检查成本高,修复失败极为困难。本文提出面向人类中心的LLM代理失败可解释性验证框架VeriLA,通过整理人类设计的代理评估标准,构建基于人类黄金标准训练的代理验证模块,系统化评估代理执行输出。该方法能以人类标准精准识别各代理的失败,提供清晰修改指引,显著降低人类认知负担。案例研究显示,VeriLA在提升可解释性和评估效率方面表现优异,有助于增强人机协作中的责任可追溯性,推动更可信、更契合人类需求的复合式AI系统发展。
原文摘要 · Abstract (English)
AI practitioners increasingly use large language model (LLM) agents in compound AI systems to solve complex reasoning tasks, these agent executions often fail to meet human standards, leading to errors that compromise the system's overall performance. Addressing these failures through human intervention is challenging due to the agents' opaque reasoning processes, misalignment with human expectations, the complexity of agent dependencies, and the high cost of manual inspection. This paper thus introduces a human-centered evaluation framework for Verifying LLM Agent failures (VeriLA), which systematically assesses agent failures to reduce human effort and make these agent failures interpretable to humans. The framework first defines clear expectations of each agent by curating human-designed agent criteria. Then, it develops a human-aligned agent verifier module, trained with human gold standards, to assess each agent's execution output. This approach enables granular evaluation of each agent's performance by revealing failures from a human standard, offering clear guidelines for revision, and reducing human cognitive load. Our case study results show that VeriLA is both interpretable and efficient in helping practitioners interact more effectively with the system. By upholding accountability in human-agent collaboration, VeriLA paves the way for more trustworthy and human-aligned compound AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。