首个系统评估AI写论文质量与幻觉风险的框架,揭示高表现伴随高幻觉。
Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers
- 用摘要重建论文再重写,分离评估呈现质量和幻觉程度。
- ClaudeCode每篇平均超10处幻觉但质量更高,Codex幻觉少但质量较低。
- 适合关注AI科研可靠性与评估标准的研究者参考。
本文提出首个系统性评估现代编码代理生成论文质量与风险的框架——Paper Reconstruction Evaluation(PaperRecon)。该框架通过从原始论文生成概要(overview.md),再由智能体基于概要和最少补充资源重建全文,并与原稿对比评估。评估分为两个正交维度:呈现质量(采用评分表)与幻觉程度(基于原稿源的代理评估)。为此构建了PaperWrite-Bench基准,包含51篇2025年后发表于顶级会议的跨领域论文。实验发现显著权衡:尽管ClaudeCode和Codex随模型演进提升,但ClaudeCode平均每篇产生超过10处幻觉,呈现质量更高;Codex幻觉更少但呈现质量较低。本工作为评估AI驱动论文写作奠定了基础,推动研究社区对相关风险的理解。
原文摘要 · Abstract (English)
This paper introduces the first systematic evaluation framework for quantifying the quality and risks of papers written by modern coding agents. While AI-driven paper writing has become a growing concern, rigorous evaluation of the quality and potential risks of AI-written papers remains limited, and a unified understanding of their reliability is still lacking. We introduce Paper Reconstruction Evaluation (PaperRecon), an evaluation framework in which an overview (overview.md) is created from an existing paper, after which an agent generates a full paper based on the overview and minimal additional resources, and the result is subsequently compared against the original paper. PaperRecon disentangles the evaluation of the AI-written papers into two orthogonal dimensions, Presentation and Hallucination, where Presentation is evaluated using a rubric and Hallucination is assessed via agentic evaluation grounded in the original paper source. For evaluation, we introduce PaperWrite-Bench, a benchmark of 51 papers from top-tier venues across diverse domains published after 2025. Our experiments reveal a clear trade-off: while both ClaudeCode and Codex improve with model advances, ClaudeCode achieves higher presentation quality at the cost of more than 10 hallucinations per paper on average, whereas Codex produces fewer hallucinations but lower presentation quality. This work takes a first step toward establishing evaluation frameworks for AI-driven paper writing and improving the understanding of its risks within the research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。