arXiv:2608.17271cs.AI2026-08被引 1

首个测试AI自主科研能力的基准,挑战模型在无指导下的创新探索。

ASI-Bench: At the Dawn of Artificial Superintelligence

论文配图:ASI-Bench: At the Dawn of Artificial Superintelligence
图 1 · 摘自论文原文
  • 设计渐进式撤回人类指导的科研任务,评估AI自主研究能力
  • 18个主流模型配置平均得分从50.91降至26.62,显示依赖人工指导
  • 覆盖11个科学领域60个项目级任务,适合推动AI自主研究发展

人工超级智能(ASI)要求AI超越已有知识的学习与应用,转向探索未知、创造新知识并转化为可验证成果。然而当前AI仍以学习、压缩和应用人类已有知识为主,现有评测多聚焦于基于已学知识的正确回答或需大量人工指导的任务完成。为此,我们推出ASI-Bench,首个联合评估AI在通用科研领域中创新探索与自主科学执行能力的基准,也是首个在同一研究项目中逐步撤回方法论指导,测试AI独立选择方法、开展研究并产出可验证结果能力的评测体系。该基准由40多位专家耗时超31,000小时构建,包含11个科学领域的60个项目级研究任务,所有任务均经专家评审、AI辅助审计、沙箱执行及评分验证。在18种前沿智能体-模型组合下,平均得分从全方法指导下的50.91下降至仅指定方法时的29.10,进一步降至自主选法时的26.62。这一显著下滑表明当前系统仍严重依赖人工指导,距离自主完成端到端项目级科研尚远。ASI-Bench向全球开放,欢迎研究者与开发者提交新任务,挑战当前AI极限,共同加速人类迈向人工超级智能的进程。官网:https://asibench.apexin.ai/submit。

原文摘要 · Abstract (English)

Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.

超级智能自主科研基准测试AI探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。