首个可复现的智能体系统行为分析数据集,揭示复杂任务下AI如何迭代决策。
Characterization of Multi-Model Agentic AI Systems on General Tasks via Trace-Driven Simulation

- 构建了双模型(MiroThinker/OWL)在GAIA上的完整执行轨迹数据集
- 通过模拟器实现低成本、可复现的系统级评估,支持多种环境配置
- 首次揭示多模型智能体在通用任务中的行为模式与设计选择的关系
智能体AI通过迭代规划、工具使用和基于观察结果的推理完成任务。尽管应用广泛,其系统级行为仍不清晰,尤其在复杂数据集和架构下,受限于高度非确定性执行、高昂评估成本及对专有模型的可见性不足。本文提出GAIATrace,首个针对两个前沿智能体系统(MiroThinker和OWL)在GAIA基准上运行的细粒度轨迹数据集,包含完整推理令牌、任务层级结构及各主要LLM的活动记录,支持深度系统研究。配套提出Vidur-Agent,一个基于轨迹驱动的模拟器,可重放GAIATrace,在多样化模拟环境中实现可复现、低成本的系统评估。利用两者,系统刻画了现代智能体处理通用任务的方式及其设计选择的影响,揭示若干独特发现。
原文摘要 · Abstract (English)
Agentic AI completes tasks through iterative planning, tool use, and reasoning based on observed outcomes. Despite its popularity, its system-level behavior remains poorly understood, particularly for complex datasets and agent architectures-owing to highly non-deterministic execution, prohibitive evaluation costs, and limited visibility into proprietary models. This paper presents GAIATrace, the first token-level trace dataset of two state-of-the-art agentic systems (MiroThinker and OWL) running GAIA, a benchmark composed of a heterogeneous mix of general-purpose tasks. Unlike prior trace datasets, GAIATrace captures full reasoning tokens, task-level structures, and activities of every major participating LLMs, enabling in-depth systems research. Complementing the dataset, we present Vidur-Agent, a trace-driven simulator that can replay GAIATrace to perform reproducible, low-cost system evaluation across diverse simulated environments. Using both artifacts, we characterize how modern agentic systems handle general tasks and how various system design choices shape their behavior, yielding several unique findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。