评测AI代理完成投行顾问复杂跨应用任务的能力,推动智能办公落地。
APEX-Agents
- 构建真实工作环境下的跨应用任务基准,模拟金融与法律行业复杂流程。
- 8个代理中最高通过率仅24.0%,显示当前AI在长周期任务上仍存明显差距。
- 开源完整数据集与执行框架,助力研究者复现与优化智能代理系统。
我们提出人工智能代理生产力指数(APEX-Agents),用于评估AI代理在投资银行分析师、管理顾问和企业律师设计的长周期、跨应用任务中的表现。该基准要求代理在包含文件和工具的真实工作环境中完成任务。我们使用Pass@1指标对八个代理进行测试,其中Gemini 3 Flash(Thinking=High)表现最佳,通过率为24.0%,其次为GPT-5.2(Thinking=High)、Claude Opus 4.5(Thinking=High)和Gemini 3 Pro(Thinking=High)。我们开源了包含480个任务的APEX-Agents基准,涵盖所有提示、评分标准、标准答案、文件及元数据。同时,我们也开源了Archipelago系统,用于代理执行与评估。
原文摘要 · Abstract (English)
We introduce the AI Productivity Index for Agents (APEX-Agents), a benchmark for assessing whether AI agents can execute long-horizon, cross-application tasks created by investment banking analysts, management consultants, and corporate lawyers. APEX-Agents requires agents to navigate realistic work environments with files and tools. We test eight agents for the leaderboard using Pass@1. Gemini 3 Flash (Thinking=High) achieves the highest score of 24.0%, followed by GPT-5.2 (Thinking=High), Claude Opus 4.5 (Thinking=High), and Gemini 3 Pro (Thinking=High). We open source the APEX-Agents benchmark (n=480) with all prompts, rubrics, gold outputs, files, and metadata. We also open source Archipelago, our infrastructure for agent execution and evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。