测试了三个大模型自动生成论文的能力,发现表面不错实则漏洞多。
How Far Are We From True Auto-Research?

- 用轻量框架让模型自主完成选题、实验、写作全流程
- 人工审核发现77%论文存在实验造假或设计缺陷
- 不同模型表现差异大,顶尖会议接受不了任何一篇
当前自动研究系统可生成完整论文,但质量与可行性不等同,该领域仍缺乏对生成论文真实水平的系统评估。我们提出ResearchArena,一个最小化框架,使现成智能体(Claude Code使用Opus 4.6,Codex使用GPT-5.4,Kimi Code使用K2.5)在轻度引导下自主完成研究全周期(构思、实验、写作、自优化)。在13个计算机科学种子任务上,每组模型-领域组合运行3次,共生成117篇论文。通过三种评估方式:仅文本评审(SAR)、含代码工作区的同行评审(PR),以及人工元评审。SAR显示,Claude Code得分最高,超过Analemma的FARS,并接近人类在ICLR 2025投稿的平均分,看似乐观;但人工检查揭示,SAR评分与实际录用决策严重脱节,仅奖励合理表述而不验证实验实质。在考虑代码工作的PR下,分数大幅下降,审计发现实验严谨性是主要瓶颈,分解为三类失败模式:虚构结果、统计效力不足、计划与执行不一致,且高度依赖模型:Codex在论文与代码不一致/伪造参考文献上的错误率分别为5%和8%,而Kimi Code高达77%和72%,差距达15倍,反映出不同模型演化出截然不同的研究人格。117篇生成论文均未达到顶级会议录用标准,表明我们离真正意义上的自动研究仍有显著距离。
原文摘要 · Abstract (English)
Recent auto-research systems can produce complete papers, but feasibility is not the same as quality, and the field still lacks a systematic study of how good agent-generated papers actually are. We introduce ResearchArena, a minimal scaffold that lets off-the-shelf agents (Claude Code using Opus 4.6, Codex using GPT-5.4, and Kimi Code using K2.5) carry out the full research loop themselves (ideation, experimentation, paper writing, self-refinement) under only lightweight guidance. Across 13 computer science seeds and 3 trials per agent-domain pair, ResearchArena yields 117 agent-generated papers, each evaluated under three complementary lenses: a manuscript-only reviewer (SAR), an artifact-aware peer review (PR) in which agents inspect the workspace alongside the manuscript, and an human conducted meta-review. Under SAR alone the picture is optimistic: Claude Code obtains the highest score, outperforms Analemma's FARS, and matches the weighted-average human ICLR 2025 submission, suggesting that minimally scaffolded agents can produce papers that look competitive on manuscript-only review. Manual inspection, however, reveals this picture is overstated: SAR scores are poorly aligned with its actual acceptance decisions and reward plausible framing without verifying experimental substance. Under artifact-aware PR scores drop sharply, and manual auditing identifies experimental rigor as the major bottleneck, decomposing into three failure modes (fabricated results, underpowered experiments, and plan/execution mismatch) that are highly agent-dependent: Codex 5%/8% paper-vs-artifact mismatch / fabricated references versus Kimi Code 77%/72%, a $\sim$15$\times$ spread that tracks distinct research personas the agents develop. None of the 117 agent-generated papers reaches the acceptance bar of a top-tier venue. This suggests that we are still gapped from the true auto-research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。