arXiv:2604.02022cs.AI2026-04被引 16

构建复杂交互下的智能体安全评估基准,揭示多阶段风险演化规律。

ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis

  • 按风险来源、失效模式、现实危害三维度设计评估框架
  • 包含1000条平均9.01轮、3.95k token的长程轨迹,覆盖1954次工具调用
  • 支持对长周期安全失效的诊断与跨模型对比,适合安全研究者使用

评估基于大语言模型的智能体安全性日益重要,因为真实部署中的风险常在多步交互中逐步显现,而非单一提示或最终响应。现有轨迹级基准受限于交互多样性不足、安全失败观测粗糙、长时序真实性弱。我们提出ATBench,一个结构化、多样化且真实的智能体轨迹安全评估基准。该基准从风险源、失效模式、现实危害三个维度组织智能体风险,构建具有异构工具池和长上下文延迟触发机制的轨迹,以捕捉多阶段的真实风险演化。基准包含1000条轨迹(503条安全,497条不安全),平均9.01轮对话,3.95千词,调用1954次工具,工具池覆盖2084种可用工具。数据质量通过规则与LLM过滤及全人工审核保障。在前沿大模型、开源模型和专用防护系统上的实验表明,即使强评估者也难以应对,同时支持分层分析、跨基准比较与长时序失效模式诊断。

原文摘要 · Abstract (English)

Evaluating the safety of LLM-based agents is increasingly important because risks in realistic deployments often emerge over multi-step interactions rather than isolated prompts or final responses. Existing trajectory-level benchmarks remain limited by insufficient interaction diversity, coarse observability of safety failures, and weak long-horizon realism. We introduce ATBench, a trajectory-level benchmark for structured, diverse, and realistic evaluation of agent safety. ATBench organizes agentic risk along three dimensions: risk source, failure mode, and real-world harm. Based on this taxonomy, we construct trajectories with heterogeneous tool pools and a long-context delayed-trigger protocol that captures realistic risk emergence across multiple stages. The benchmark contains 1,000 trajectories (503 safe and 497 unsafe), averaging 9.01 turns and 3.95k tokens, with 1,954 invoked tools drawn from pools spanning 2,084 available tools. Data quality is supported by rule-based and LLM-based filtering plus full human audit. Experiments on frontier LLMs, open-source models, and specialized guard systems show that ATBench is challenging even for strong evaluators, while enabling taxonomy-stratified analysis, cross-benchmark comparison, and diagnosis of long-horizon failure patterns.

智能体安全轨迹评估风险诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。