arXiv:2604.24697cs.AI2026-04被引 1

用红石电路测试AI从发现到应用的闭环能力,发现顶级模型仅26%成功。

Can Current Agents Close the Discovery-to-Application Gap? A Case Study in Minecraft

论文配图:Can Current Agents Close the Discovery-to-Application Gap? A Case Study in Minecraft
图 1 · 摘自论文原文
  • 在Minecraft中设计可扩展红石电路任务,逼迫AI真正发现规律而非记忆答案。
  • 所有前沿模型成功率均卡在约26%,无法完成复杂电路搭建。
  • 揭示当前瓶颈从“解决问题”转向“提出正确问题”,适合研究AI认知局限的学者。

发现因果规律并将其应用于构建功能性系统——即发现到应用的循环——是通用智能的核心特征,但科学发现与现实工程之间的巨大复杂性鸿沟阻碍了对这一能力的评估。我们提出SciCrafter,一个基于Minecraft的基准,通过参数化红石电路任务来实现该循环。代理需按指定模式(如同时或定时序列)点亮灯泡;目标参数的增加显著提升构建复杂度和所需知识量,迫使真正的发现而非依赖记忆解法。在通用代码代理框架下评估前沿模型GPT-5.2、Gemini-3-Pro和Claude-Opus-4.5,结果发现所有模型的成功率均停滞在约26%。为诊断失败原因,我们将循环分解为四种能力:知识缺口识别、实验发现、知识整合与知识应用,并设计针对性干预措施,其边际贡献作为对应差距的代理指标。分析显示,尽管通用知识应用仍是所有模型的最大短板,但对于前沿模型而言,知识缺口识别已开始成为主要障碍,表明瓶颈正从‘解决难题’转向‘提出正确问题’。我们发布SciCrafter作为未来研究人工智能系统在完整发现-应用循环中导航能力的诊断工具。

原文摘要 · Abstract (English)

Discovering causal regularities and applying them to build functional systems--the discovery-to-application loop--is a hallmark of general intelligence, yet evaluating this capacity has been hindered by the vast complexity gap between scientific discovery and real-world engineering. We introduce SciCrafter, a Minecraft-based benchmark that operationalizes this loop through parameterized redstone circuit tasks. Agents must ignite lamps in specified patterns (e.g., simultaneously or in timed sequences); scaling target parameters substantially increases construction complexity and required knowledge, forcing genuine discovery rather than reliance on memorized solutions. Evaluating frontier models including GPT-5.2, Gemini-3-Pro, and Claude-Opus-4.5 under a general-purpose code agent scaffold, we find that all plateau at approximately 26% success rate. To diagnose these failures, we decompose the loop into four capacities--knowledge gap identification, experimental discovery, knowledge consolidation, and knowledge application--and design targeted interventions whose marginal contributions serve as proxies for corresponding gaps. Our analysis reveals that although the general knowledge application capability still remains as the biggest gap across all models, for frontier models the knowledge gap identification starts to become a major hurdle--indicating the bottleneck is shifting from solving problems right to raising the right problems for current AI. We release SciCrafter as a diagnostic probe for future research on AI systems that navigate the full discovery-to-application loop.

AI评估红石电路发现-应用智能评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。