arXiv:2511.08042cs.AI2025-11被引 2

为企用智能体设计新评估标准,打破旧榜单迷思

Towards a Standard, Enterprise-Relevant Agentic AI Benchmark: Lessons from 5.5 billion tokens' worth of agentic AI evaluations

  • 构建抗数据污染的企用智能体评测基准KAMI v0.1
  • 5.5亿令牌测试显示新模型未必优于旧版
  • 适合关注部署成本与推理效率的企业决策者

企业采用智能体系统需可靠评估方法以反映真实部署场景。传统大模型基准存在训练数据污染问题,且无法衡量多步工具使用、不确定环境下的决策等智能体能力。我们提出面向企业的Kamiwaza智能体价值指数(KAMI)v0.1,解决污染抗性和智能体评估问题。通过35种模型配置、17万条测试题处理超过55亿令牌,结果表明传统基准排名难以预测实际智能体表现。值得注意的是,如Llama 4或Qwen 3等新一代模型在企业相关任务中并不总优于其旧版,与传统基准趋势相悖。我们还揭示了成本-性能权衡、模型行为模式差异及推理能力对令牌效率的影响,这些发现对企业部署决策至关重要。

原文摘要 · Abstract (English)

Enterprise adoption of agentic AI systems requires reliable evaluation methods that reflect real-world deployment scenarios. Traditional LLM benchmarks suffer from training data contamination and fail to measure agentic capabilities such as multi-step tool use and decision-making under uncertainty. We present the Kamiwaza Agentic Merit Index (KAMI) v0.1, an enterprise-focused benchmark that addresses both contamination resistance and agentic evaluation. Through 170,000 LLM test items processing over 5.5 billion tokens across 35 model configurations, we demonstrate that traditional benchmark rankings poorly predict practical agentic performance. Notably, newer generation models like Llama 4 or Qwen 3 do not always outperform their older generation variants on enterprise-relevant tasks, contradicting traditional benchmark trends. We also present insights on cost-performance tradeoffs, model-specific behavioral patterns, and the impact of reasoning capabilities on token efficiency -- findings critical for enterprises making deployment decisions.

智能体评估企业应用大模型评测推理效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。