arXiv:2605.08545cs.AI2026-05被引 3

仅看通过/失败结果会误导评估,必须分析日志才能真实反映AI代理能力。

Log analysis is necessary for credible evaluation of AI agents

  • 通过系统记录和分析代理的输入、执行与输出来验证评估可信度。
  • 在tau-Bench Airline中发现通过率被低估近50%,且暴露了隐藏的部署缺陷。
  • 为评测者、开发者和部署方提供可落地的日志分析指南,提升评估透明度。

当前智能体评测通常仅报告最终结果:成功或失败。这会带来三重可信度威胁:一是评分可能因捷径和基准漏洞被夸大或低估,扭曲真实能力;二是评测表现可能无法预测实际应用价值,因评测框架存在局限性及重复失败模式;三是能力分数可能掩盖代理采取的危险甚至灾难性行为。本文主张,日志分析——即对智能体的输入、执行过程与输出进行系统追踪与分析——是克服这些有效性威胁、实现可信评测的必要手段。本文提出一个基于日志分析的可信评测威胁分类体系,并建立一套指导原则。以tau-Bench Airline为例,揭示其通过率实际被低估近50%,并发现仅凭结果指标无法察觉的部署失败模式。最后,针对评测设计者、模型开发者、独立评估者和部署方,提出切实可行的推广建议。

原文摘要 · Abstract (English)

Agent benchmarks typically report only final outcomes: pass or fail. This threatens evaluation credibility in three ways. First, scores may be inflated or deflated by shortcuts and benchmark artifacts, misrepresenting capability. Second, benchmark performance may fail to predict real-world utility due to scaffold limitations and recurring failure modes. Finally, capability scores may conceal dangerous or catastrophic actions taken by the agent. We argue that log analysis -- the systematic tracking and analysis of the inputs, execution, and outputs of an AI agent -- is necessary to overcome these validity threats and promote credible agent evaluation. In this paper, we (1) present a taxonomy of threats to credible evaluation documented through log analysis, and (2) develop a set of guiding principles for log analysis. We illustrate these principles on tau-Bench Airline, revealing that pass^5 performance was under-elicited by nearly 50% and surfacing deployment failure modes invisible to outcome metrics. We conclude with pragmatic recommendations to increase uptake of log analysis, directed at diverse stakeholders including benchmark creators, model developers, independent evaluators, and deployers.

AI评估日志分析智能体评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。