arXiv:2604.03338econ.GNcs.AI2026-04被引 1

AI论文在选题上远逊于人类,执行能力次之,选题是主要瓶颈。

The Ideation Bottleneck: Decomposing the Quality Gap Between AI-Generated and Human Economics Research

  • 用双模型评估选题与执行质量,方法科学且一致。
  • 人类论文选题优秀率47.1%,AI仅16.5%;执行得分人类4.38,AI3.84。
  • 74% AI论文用同一方法,仅0.8%同时超越人类平均水平。

自主AI系统已能生成完整经济学论文,但在与人类论文的对比中表现显著落后。本文将质量差距分解为选题质量和执行质量两个独立维度。采用微调语言模型集成评估选题质量(基于Gong、Li、Zhou, 2026数据),并使用与APE竞赛裁判同源的Gemini 3.1 Flash Lite,通过六维度量表评估执行质量,分析953篇论文(912篇AI生成自APE项目,41篇人类发表于AER和AEJ: Economic Policy)。选题差距巨大(Cohen's d = 2.23, p < 0.001),人类平均优秀概率47.1%,AI仅为16.5%;执行差距显著但较小(d = 0.90, p < 0.001),人类得分4.38/5.0,AI为3.84。选题贡献了约71%的整体差距,执行占29%。执行中最弱环节是机制分析深度(d = 1.43);稳健性无显著差异。74%的AI论文使用差分法,仅7篇(0.8%)在选题与执行上均超过人类中位水平。当前竞争性经济学研究的主要瓶颈仍在于选题创新。

原文摘要 · Abstract (English)

Autonomous AI systems can now generate complete economics research papers, but they substantially underperform human-authored publications in head-to-head comparisons. This paper decomposes the quality gap into two independent components: research idea quality and execution quality. Using a two-model ensemble of fine-tuned language models trained on publication decisions (Gong, Li, and Zhou, 2026) to evaluate idea quality and a comprehensive six-dimension rubric assessed by Gemini 3.1 Flash Lite -- the same model family used as the APE tournament judge, ensuring methodological consistency -- to evaluate execution quality, we analyze 953 economics papers -- 912 AI-generated papers from the APE project and 41 human papers published in the American Economic Review and AEJ: Economic Policy. The idea quality gap is large (Cohen's d = 2.23, p < 0.001), with human papers achieving 47.1% mean ensemble exceptional probability versus 16.5% for AI. The execution quality gap is also significant but smaller (d = 0.90, p < 0.001), with human papers scoring 4.38/5.0 versus 3.84. Idea quality accounts for approximately 71% of the overall quality difference, with execution contributing 29%. The largest execution weakness is mechanism analysis depth (d = 1.43); no significant difference is found on robustness. We document that 74% of AI papers employ difference-in-differences, and only 7 AI papers (0.8%) surpass the median human paper on both idea and execution quality simultaneously. The primary bottleneck to competitive AI-generated economics research remains ideation.

AI科研选题质量经济学评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。