首个统一评估各类智能体的基准,涵盖从强化学习到大模型的多种方法。
Agentick: A Unified Benchmark for General Sequential Decision-Making Agents

- 构建37个程序生成任务,支持多模态输入与统一接口。
- 大模型推理增强使性能提升3-10倍,文本观察优于自然语言。
- 适合研究通用决策智能体的学者和开发者使用。
AI智能体研究涵盖从强化学习到基础模型的多种范式,但缺乏统一基准进行公平比较。我们提出Agentick,一个面向序列决策智能体的基准,可评估强化学习、大语言模型、视觉语言模型、混合及人类智能体。该基准包含37个程序生成任务,覆盖六类能力、四个难度等级和五种观察模态,通过单一Gymnasium兼容接口提供。内置编码API、所有任务的参考策略、预构建SFT数据集、可组合的智能体框架及实时排行榜。对27种配置超过9万次试验的评估显示:无单一方法全面领先;GPT-5 mini总体得分最高(0.309归一化分);PPO在规划与多智能体任务中占优;推理增强使大模型性能提升3-10倍;ASCII观测始终优于自然语言。结果表明各类智能体仍有巨大改进空间。Agentick的分解式、多模态设计为推动通用自主智能体的发展提供了实证基础设施,既可用作评估框架,也可作为基础模型强化学习后训练的真实序列环境。
原文摘要 · Abstract (English)
AI agent research spans a wide spectrum: from RL agents that learn from scratch to foundation model agents that leverage pre-trained knowledge, yet no unified benchmark enables fair comparison across these approaches. We present Agentick, a benchmark for sequential decision-making agents designed to evaluate RL, LLM, VLM, hybrid, and human agents on common ground and to power research on the fundamental challenges of sequential decision-making. Agentick provides 37 procedurally generated tasks across six capability categories, four difficulty levels, and five observation modalities, all exposed through a single Gymnasium-compatible interface. The benchmark ships with a Coding API, oracle reference policies for all tasks, pre-built SFT datasets, a composable agent harness, and a live leaderboard. An evaluation spanning 27 configurations and over 90,000 episodes reveals that no single approach dominates: GPT-5 mini leads overall at 0.309 oracle-normalized score while PPO dominates planning and multi-agent tasks; the reasoning harness multiplies LLM performance by 3-10x; and ASCII observations consistently outperform natural language. These findings highlight the substantial room for improvement that remains across all agent paradigms. Agentick's capability-decomposed, multi-modal design provides the empirical infrastructure needed to drive progress toward general autonomous agents, both as an evaluation framework and as a training ground for RL post-training of foundation models in truly sequential environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。