arXiv:2507.05241cs.AIcs.CL2025-07被引 46

构建可自主科研的通用AI agent,首次在人类终极考题中突破30%准确率

SciMaster: Towards General-Purpose Scientific AI Agents, Part I. X-Master as Foundation: Can We Lead on Humanity's Last Exam?

  • 设计工具增强型推理架构,用代码作为交互语言灵活调用工具与库
  • 在人类终极考题上达32.1%正确率,首次超越30%阈值,领先OpenAI与谷歌
  • 开源方案可复现,为未来通用科学智能提供训练与评估范式

AI代理的快速发展激发了加速科学发现的长期愿景。实现这一目标需深刻理解人类知识前沿。为此,人类终极考题(HLE)成为评估科学AI代理的严苛基准。本文旨在构建通用代理的基础架构,并通过在HLE上取得领先表现来验证其能力。我们提出X-Master,一种工具增强型推理代理,通过灵活调用外部工具模拟人类研究者行为。该代理以代码为交互语言,可动态使用内置Python库及自定义工具扩展推理能力。进一步通过X-Masters——一种分散堆叠的智能体工作流,系统性提升推理广度与深度。我们的开源方案在HLE上达到32.1%准确率,超越OpenAI(26.6%)和Google Deep Research(26.9%),首次突破30%门槛。本工作深化了对复杂任务求解的理解,积累了宝贵经验,可指导后续模型训练。

原文摘要 · Abstract (English)

The rapid advancements of AI agents have ignited the long-held ambition of leveraging them to accelerate scientific discovery. Achieving this goal requires a deep understanding of the frontiers of human knowledge. As such, Humanity's Last Exam (HLE) provides an exceptionally challenging touchstone for evaluating scientific AI agents. In this work, we aim to construct the foundational architecture for general-purpose agents and validate the capabilities through leading performance on HLE. To achieve this, we introduce X-Master, a tool-augmented reasoning agent designed to emulate human researchers by interacting flexibly with external tools during its reasoning process. This agent, guided by the conceptualization of code as an interaction language, can flexibly leverage built-in Python libraries and our customized tools to augment the reasoning. We further scale its capabilities through X-Masters, a scattered-and-stacked agentic workflow that systematically enhances breadth and depth of reasoning. Our open-source solution, X-Masters, sets a new state-of-the-art record on HLE with a score of 32.1%, surpassing OpenAI's and Google's Deep Research (26.6% and 26.9%) and becoming the first to exceed the 30% threshold. This work allows us to gain a deeper understanding of complex task-solving and accumulates valuable experience that can inform future advancements, guiding subsequent model training.

科学AI智能体通用智能推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。