arXiv:2606.07462cs.AI2026-06被引 1

测试大模型能否像真实研究员一样完成科研全流程任务。

Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle

论文配图:Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle
图 1 · 摘自论文原文
  • 设计新基准AARRI-Bench,评估代理在科研细节中的表现
  • 顶尖模型成功率仅68.3%,常忽略关键细节
  • 适合关注AI科研助手可信性与局限性的研究者

随着基础模型发展和智能体架构日益复杂,智能体已在长周期编码任务甚至自主实验中展现卓越能力。然而,尽管其已从研究助手演变为自主研究智能体,仍存在领域敏感性不足、科研伦理缺失及科学判断力薄弱等显著局限,难以完全替代人类研究者。为此,我们提出AARR(Act As a Real Researcher)基准系列,聚焦于智能体是否能模拟真实研究者在微观科研场景中的专业性、严谨性和精细推理。本文推出该系列首个基准AARRI-Bench(Act As a Real Research Intern)。我们在前沿模型与智能体系统上开展广泛实验,发现即使最佳配置(Mini-SWE-Agent + Claude Opus 4.7)也仅达68.3%的成功率,且频繁遗漏对人类研究者而言显而易见的细微但关键信息。结果表明,构建类研究员的AI需深入探索研究行为本质,而非仅依赖复杂架构。数据已开源:https://github.com/AARR-bench/AARRI-bench。

原文摘要 · Abstract (English)

As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and even autonomous experiment execution. Despite their evolution from research assistants into autonomous research agents, these systems still exhibit significant limitations in field sensitivity, research ethics, and nuanced scientific judgment. Consequently, frontier agents remain unable to fully replace human researchers. To bridge this gap, we conceptualize the AARR (Act As a Real Researcher) benchmark series. Unlike existing benchmarks that primarily assess macro-level execution capabilities, AARR focuses on whether agents can emulate the professionalism, thoroughness, and nuanced reasoning that characterize human researchers in granular research scenarios. In this work, we propose AARRI-Bench (Act As a Real Research Intern), the first benchmark in this series. We conduct extensive experiments across frontier models and agentic systems, revealing that even the best-performing configuration (Mini-SWE-Agent with Claude Opus 4.7) achieves only 68.3\% success rate, frequently overlooking subtle yet critical details that are obvious to real human researchers. Our results indicate that developing researcher-like AI requires further exploration of research behavior, rather than merely complex scaffolding. Our data is released at https://github.com/AARR-bench/AARRI-bench.

智能体科研自动化基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。