arXiv:2605.16282cs.CYcs.AI2026-05

首次系统分析智能体安全评测基准,揭示评测方法混乱与结论不可靠问题。

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents

  • 构建六轴分类体系,梳理40个安全评测基准的方法论差异。
  • 发现95%置信区间下评测排名无一致性,同一模型在不同基准表现矛盾。
  • 指出当前评测过度关注外部风险,忽略智能体内部潜在危险,且缺乏统一标准。

基于大语言模型的自主智能体快速部署带来了远超传统大模型的安全风险,自2023年底以来催生了大量安全评测基准。然而这些基准独立发展,威胁模型不一致、评估指标不兼容、风险覆盖重叠且不完整。本文首次系统分析作为评估工具的智能体安全基准,梳理了2023至2026年间40个行为型智能体安全评测基准,以及5个相关评估器、防御方案和数据集工具。提出六轴评测方法分类体系,并应用于全集,揭示方法选择如何影响安全结论。覆盖矩阵显示风险覆盖面广但方法论分歧显著;分类分析表明核心评测集中于沙盒化、受限环境且多为纯安全评估。跨基准一致性检验(95%置信区间,Kendall's W分析)显示,各维度评分无排名一致性(W = 0.10, p = 0.94),基准选择可导致矛盾安全结论,覆盖率常夸大评估深度,环境保真度系统性影响报告安全性,领域偏重外部施加风险而非智能体内在风险,指标碎片化阻碍比较,鲁棒性基本未被评测。研究释放结构化元数据、完整分类编码、风险标注及所有实验资料,并提出未来评测的最低报告标准。

原文摘要 · Abstract (English)

The rapid deployment of LLM-based autonomous agents has introduced safety risks that extend far beyond traditional LLM concerns, prompting a proliferation of safety benchmarks since late 2023. However, these benchmarks have developed independently, with inconsistent threat models, incompatible metrics, and overlapping yet incomplete risk coverage. We present the first systematic analysis dedicated to agent safety benchmarks as evaluation instruments. We catalog 40 behavioral agent-safety benchmarks (2023-2026), plus 5 adjacent evaluator, defense, and dataset artifacts, propose a six-axis taxonomy of benchmark evaluation methodology, and apply it across the corpus to characterize how methodological choices shape safety conclusions. A coverage matrix reveals broad risk coverage but limited methodological convergence, while the taxonomy analysis shows a behavioral-benchmark core concentrated in sandboxed, constrained, and often safety-only evaluation. Across the landscape, we find that benchmark choice can yield contradictory safety conclusions, coverage counts often overstate evaluation depth, environment fidelity systematically shapes reported safety, the field disproportionately tests externally imposed rather than agent-internal risks, metric fragmentation limits comparison, and robustness remains effectively unbenchmarked. We ground these claims with a cross-benchmark consistency check, with 95% confidence intervals and Kendall's W concordance analysis, finding no evidence of ranking concordance across evaluation dimensions (W = 0.10, p = 0.94). We release structured metadata, full taxonomy codings, risk annotations, and all experimental artifacts, and propose minimum reporting standards for future benchmarks.

安全评测智能体基准分析大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。