arXiv:2605.22568cs.CRcs.AI2026-05被引 2
安全类AI代理评估难,三大痛点需破解
Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard
- 指出安全评估基准的三大缺陷:漏洞、过时、运行不确定性
- 实证揭示现有评测方法易被误导,结果不可靠
- 为构建可信评估体系提供可操作改进方向
用于评估安全关键角色中AI代理的基准测试存在严重缺陷。基于近期实证证据,我们识别出三项核心挑战:基准漏洞、时间过时性以及运行时不确定性。这些问题导致安全评估结果失真,难以真实反映代理能力。本文进一步提出切实可行的方向,以构建更稳健、可信的评估框架,推动安全代理评价体系的实质性改进。
原文摘要 · Abstract (English)
The benchmarks used to evaluate AI agents in security-critical roles suffer from crucial weaknesses. Building on recent empirical evidence, we characterize three core challenges that undermine security evaluations: benchmark vulnerabilities, temporal staleness, and runtime uncertainty. We then outline practical directions toward building more robust and trustworthy evaluation frameworks.
AI评估安全评测基准测试
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。