arXiv:2510.05432cs.AI2025-10

大模型仅靠记忆知识就能解决科研问题吗?

AInstein: Can LLMs Solve Research Problems From Parametric Memory Alone?

  • 用迭代批判循环让大模型自主生成并优化科研解法
  • 70%以上问题能成功解决,但仅19%能复现原文方法
  • 适合研究大模型自主科研能力的边界与潜力

大语言模型能否在不进行微调、无外部检索的情况下,仅凭参数化知识解决人工智能研究问题?我们提出AInstein框架,通过迭代批判循环测试大模型生成和改进科研解决方案的能力。对20位领域专家进行盲测,验证了自动化评估指标的有效性,随后将该指标扩展至1,214篇ICLR 2025论文,采用大模型作为评判者。两个核心指标分别衡量:成功率(方案是否解决问题)和重发现率(是否匹配已发表方法)。结果显示,大模型在超过70%的问题上取得成功,但严格复现原文方法的比例不足19%,表明其具备真正的推理能力而非简单回忆。然而,该能力存在明显边界:模型在熟悉的方法论领域表现良好,但在需要跨领域类比迁移时失败,这种现象称为参数知识边界。在ResearchPlanGen基准(2,645个问题)上,无需训练的迭代优化策略达到强化学习微调水平;通过准则覆盖分析,进一步明确了测试时精炼所能达到的上限。这些发现共同刻画了大模型作为自主科研求解者的潜力与局限。

原文摘要 · Abstract (English)

Can large language models solve AI research problems using only their parametric knowledge, without fine-tuning, retrieval, or other external aids? We introduce AInstein, a framework for testing whether LLM agents can generate and refine solutions to research problems through iterative critique loops. A blind study with 20 domain experts on held-out ICLR 2026 problems validates our automated metrics, which we then scale to 1,214 ICLR 2025 papers using an LLM-as-a-judge paradigm. Two metrics capture complementary aspects of performance: Success Rate (does the solution address the problem?) and Rediscovery (does it match the published approach?). LLMs succeed on over 70% of problems, yet strictly rediscover the published solution less than 19% of the time, suggesting genuine problem-solving rather than associative recall. However, this ability has clear limits: models handle familiar methodological territory well but fail when solutions require cross-domain analogical transfer, a pattern we call the parametric knowledge boundary. On the ResearchPlanGen benchmark (2,645 problems), our training-free iterative refinement strategy matches RL finetuning, and a criteria-coverage analysis pins down the ceiling of what test-time refinement alone can achieve. Together, these findings map both the capabilities and the limits of LLMs as autonomous scientific problem-solvers.

大模型科研自动化自主求解能力边界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。