arXiv:2511.14136cs.AI2025-11被引 5

提出多维度评估框架,解决企业AI代理系统成本与可靠性被忽视的问题。

Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems

  • 构建涵盖成本、延迟、可靠性等五维的CLEAR评估框架
  • 实测显示仅追求准确率的代理成本高出4.4至10.8倍
  • 适合关注生产落地的AI架构师与企业决策者

当前代理型AI基准测试主要评估任务完成准确率,忽视企业部署中至关重要的成本效益、可靠性与运行稳定性。通过对12个主流基准的系统分析及对先进代理的实证评估,我们发现三大根本缺陷:(1)缺乏成本控制评估,相同精度下成本差异达50倍;(2)可靠性评估不足,代理性能从单次运行60%下降至八次一致性测试仅25%;(3)缺少安全、延迟与合规性等多维指标。为此,我们提出针对企业部署设计的综合评估框架CLEAR(Cost, Latency, Efficacy, Assurance, Reliability)。在300个企业任务上对六款领先代理的评估表明,仅优化准确率的代理比成本感知型方案贵4.4至10.8倍,且性能相当。专家评估(N=15)证实,CLEAR对生产成功预测的相关性为ρ=0.83,显著高于仅看准确率的ρ=0.41。

原文摘要 · Abstract (English)

Current agentic AI benchmarks predominantly evaluate task completion accuracy, while overlooking critical enterprise requirements such as cost-efficiency, reliability, and operational stability. Through systematic analysis of 12 main benchmarks and empirical evaluation of state-of-the-art agents, we identify three fundamental limitations: (1) absence of cost-controlled evaluation leading to 50x cost variations for similar precision, (2) inadequate reliability assessment where agent performance drops from 60\% (single run) to 25\% (8-run consistency), and (3) missing multidimensional metrics for security, latency, and policy compliance. We propose \textbf{CLEAR} (Cost, Latency, Efficacy, Assurance, Reliability), a holistic evaluation framework specifically designed for enterprise deployment. Evaluation of six leading agents on 300 enterprise tasks demonstrates that optimizing for accuracy alone yields agents 4.4-10.8x more expensive than cost-aware alternatives with comparable performance. Expert evaluation (N=15) confirms that CLEAR better predicts production success (correlation $ρ=0.83$) compared to accuracy-only evaluation ($ρ=0.41$).

AI评估企业应用多维指标成本优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。