四次自主生成论文尝试中仅一次成功,揭示大模型做科研的六大瓶颈。
Why LLMs Aren't Scientists Yet: Lessons from Four Autonomous Research Attempts
- 用六个LLM代理模拟科研全流程,自动完成从选题到投稿。
- 三例失败于实现或验证阶段,一例被实验会议接收并通过多轮评审。
- 暴露模型偏见、执行漂移、记忆衰退等缺陷,需强化科学判断力。
我们报告了四次端到端自主生成机器学习研究论文的案例研究,采用六名LLM代理映射至科研流程的各个阶段。其中三次在实施或评估阶段失败,一次完整通过流程并被Agents4Science 2025录用——该会议允许AI系统作为第一作者,且通过了人工与多智能体评审。研究总结出六种反复出现的失败模式:对训练数据默认值的偏好、执行压力下的实现漂移、长时任务中的记忆与上下文退化、明知失败仍宣称成功的过度兴奋、领域知识不足,以及实验设计中的科学品味薄弱。最后提出四项设计原则,讨论其对自主科学发现的意义,并公开所有提示词、成果及输出文件于https://github.com/Lossfunk/ai-scientist-artefacts-v1。
原文摘要 · Abstract (English)
We report a case study of four end-to-end attempts to autonomously generate ML research papers using a pipeline of six LLM agents mapped to stages of the scientific workflow. Of these four, three attempts failed during implementation or evaluation. One completed the pipeline and was accepted to Agents4Science 2025, an experimental inaugural venue that required AI systems as first authors, passing both human and multi-AI review. From these attempts, we document six recurring failure modes: bias toward training data defaults, implementation drift under execution pressure, memory and context degradation across long-horizon tasks, overexcitement that declares success despite obvious failures, insufficient domain intelligence, and weak scientific taste in experimental design. We conclude by discussing four design principles for more robust AI-scientist systems, implications for autonomous scientific discovery, and we release all prompts, artifacts, and outputs at https://github.com/Lossfunk/ai-scientist-artefacts-v1
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。