为长时序智能体行为提供细粒度归因分析的统一基准与标注框架
Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework
- 构建统一组件架构,对智能体轨迹进行细粒度标注
- 涵盖1300+条轨迹,覆盖任务动作、危险行为与拒绝响应
- 支持局部、长程及结构化归因链分析,适合安全评估研究者
大型语言模型智能体在执行长时序任务时涉及用户指令、工具调用、外部观察和记忆等复杂交互。现有基准多关注行为结果,缺乏对行为归因的细粒度支持。本文提出轨迹归因任务,构建统一组件架构的基准与标注框架,包含主归因组件、攻击链与执行链标注。基于AgentDojo及Agent3Sigma的Stage与Canary设置,共收集超过1300条标注轨迹,涵盖任务对齐动作、不安全行为与安全拒绝。定义两项评估任务:主归因定位与归因链恢复,并提供基于增量贡献与组件级留一扰动的基线方法。结果揭示不同归因场景下性能差异显著,反映该基准的挑战性。此外,发布可复用的标注技能,支持新智能体轨迹的标准化标注与评估。项目资源与后续更新见https://github.com/chenjing-2024/agent-trajectory-attribution。
原文摘要 · Abstract (English)
Large language model (LLM) agents increasingly operate through long-horizon trajectories involving user instructions, tool use, external observations, and memory. Existing benchmarks primarily evaluate behavioral outcomes but provide limited support for fine-grained attribution analysis. We introduce trajectory attribution and develop a benchmark and annotation framework for this task. The benchmark organizes heterogeneous trajectories under a unified component schema and provides annotations of the primary attribution component, together with attack and execution chains where applicable. Instantiating the benchmark with trajectories from AgentDojo and the Stage and Canary settings of Agent3Sigma yields more than 1,300 annotated trajectories covering task-aligned actions, unsafe actions, and safety refusals. The benchmark defines two evaluation tasks, primary attribution localization and attribution-chain recovery, and provides reference baselines based on incremental trajectory contribution and component-level leave-one-out perturbation. It captures diverse attribution settings, including local and long-range attribution as well as structured attribution chains. Reference baseline results exhibit substantial performance differences across these settings, providing an initial characterization of the benchmark's attribution challenges. Beyond this initial instantiation, we release a reusable annotation skill that enables trajectories generated by new agent models to be standardized, annotated, and evaluated under the same framework. Project resources and future releases are available at https://github.com/chenjing-2024/agent-trajectory-attribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。