arXiv:2605.02661cs.AIcs.CY2026-05

面向学生真实学术挑战的多领域智能体评测基准。

AcademiClaw: When Students Set Challenges for AI Agents

论文配图:AcademiClaw: When Students Set Challenges for AI Agents
图 1 · 摘自论文原文
  • 从学生真实学术任务中提取80个复杂长周期问题
  • 顶尖模型仅55%通过率,揭示能力边界与行为差异
  • 适合评估智能体在科研、编程等真实场景中的表现

OpenClaw生态中的基准测试长期仅限于助理级任务,未充分评估其学术级能力。我们提出AcademiClaw,一个由80个复杂长周期任务构成的双语基准,源自大学生真实学术工作流——作业、研究项目、竞赛和自研项目,这些任务均是当前AI代理难以有效解决的。任务经230个学生提交候选的严格专家评审筛选,涵盖25个以上专业领域,从奥数级数学与语言学问题到需CUDA GPU执行的强化学习与全栈系统调试,其中16项需GPU运行。每项任务在隔离Docker沙箱中执行,采用六种互补评分技术结合多维评分标准进行完成度评估,并有独立五类安全审计提供行为分析。六款前沿模型实验显示,即使最优者也仅达55%通过率。进一步分析揭示了任务领域间显著的能力分界、模型间策略差异,以及令牌消耗与输出质量之间的脱节,提供了超越聚合指标的细粒度诊断信号。我们希望AcademiClaw及其开源数据与代码能为OpenClaw社区提供实用资源,推动更强大、更通用的智能体发展,以应对真实世界学术需求。所有数据与代码见https://github.com/GAIR-NLP/AcademiClaw。

原文摘要 · Abstract (English)

Benchmarks within the OpenClaw ecosystem have thus far evaluated exclusively assistant-level tasks, leaving the academic-level capabilities of OpenClaw largely unexamined. We introduce AcademiClaw, a bilingual benchmark of 80 complex, long-horizon tasks sourced directly from university students' real academic workflows -- homework, research projects, competitions, and personal projects -- that they found current AI agents unable to solve effectively. Curated from 230 student-submitted candidates through rigorous expert review, the final task set spans 25+ professional domains, ranging from olympiad-level mathematics and linguistics problems to GPU-intensive reinforcement learning and full-stack system debugging, with 16 tasks requiring CUDA GPU execution. Each task executes in an isolated Docker sandbox and is scored on task completion by multi-dimensional rubrics combining six complementary techniques, with an independent five-category safety audit providing additional behavioral analysis. Experiments on six frontier models show that even the best achieves only a 55\% pass rate. Further analysis uncovers sharp capability boundaries across task domains, divergent behavioral strategies among models, and a disconnect between token consumption and output quality, providing fine-grained diagnostic signals beyond what aggregate metrics reveal. We hope that AcademiClaw and its open-sourced data and code can serve as a useful resource for the OpenClaw community, driving progress toward agents that are more capable and versatile across the full breadth of real-world academic demands. All data and code are available at https://github.com/GAIR-NLP/AcademiClaw.

AI评测学术智能体长程任务多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。