构建可验证的多智能体评估框架,提升企业级AI系统的可信与可控性。
AEMA: Verifiable Evaluation Framework for Trustworthy and Controlled Agentic LLM Systems
- 设计流程感知的多步评估机制,支持异构智能体工作流的自动化执行
- 相比单一模型评分,评估结果更稳定、与人类判断更一致
- 适合需要透明审计的企业级智能体系统开发与部署
基于大语言模型的多智能体系统评估仍面临重大挑战,这类系统需在动态任务中表现出可靠的协作能力、透明的决策过程和可验证的性能。现有评估方法多局限于单次响应评分或狭窄基准,难以在企业级多智能体场景下实现稳定、可扩展与自动化。本文提出AEMA(自适应多智能体评估),一个流程感知且可审计的框架,可在人工监督下规划、执行并聚合跨异构智能体工作流的多步骤评估。相较于单一语言模型作为裁判,AEMA在稳定性、人类对齐性和可追溯记录方面表现更优。在基于真实业务场景模拟的企业级智能体工作流测试中,AEMA展现出透明、可复现的负责任评估路径。
原文摘要 · Abstract (English)
Evaluating large language model (LLM)-based multi-agent systems remains a critical challenge, as these systems must exhibit reliable coordination, transparent decision-making, and verifiable performance across evolving tasks. Existing evaluation approaches often limit themselves to single-response scoring or narrow benchmarks, which lack stability, extensibility, and automation when deployed in enterprise settings at multi-agent scale. We present AEMA (Adaptive Evaluation Multi-Agent), a process-aware and auditable framework that plans, executes, and aggregates multi-step evaluations across heterogeneous agentic workflows under human oversight. Compared to a single LLM-as-a-Judge, AEMA achieves greater stability, human alignment, and traceable records that support accountable automation. Our results on enterprise-style agent workflows simulated using realistic business scenarios demonstrate that AEMA provides a transparent and reproducible pathway toward responsible evaluation of LLM-based multi-agent systems. Keywords Agentic AI, Multi-Agent Systems, Trustworthy AI, Verifiable Evaluation, Human Oversight
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。