arXiv:2603.15164cs.CLcs.AI2026-03被引 5

用未来论文影响力评估AI生成的研究想法,发现大模型评价常高估不靠谱的新点子。

HindSight: Evaluating LLM-Generated Research Ideas via Future Impact

  • 通过对比生成想法与30个月后真实发表论文的引用和录用情况来评估质量。
  • 检索增强型方法生成的想法在真实研究中得分高出2.5倍,而大模型评价看不出差异。
  • 大模型评出的‘新颖’想法反而更难落地,适合想验证想法真实价值的研究者。

评估AI生成的研究想法通常依赖大模型评判或人工评审——两者主观且脱离实际研究影响。本文提出HindSight,一种基于时间切分的评估框架:将想法生成系统限定于T时刻前的文献,再评估其输出在后续30个月内对应的真实发表论文的引用量与会议接受率。在10个AI/ML研究主题上实验显示,大模型评判认为检索增强与普通生成无显著差异(p=0.584),但HindSight表明检索增强系统生成的想法评分高出2.5倍(p<0.001)。此外,大模型评价的新颖性与真实影响力呈负相关(ρ=-0.29,p<0.01),说明大模型系统性高估了难以落地的“新奇”点子。

原文摘要 · Abstract (English)

Evaluating AI-generated research ideas typically relies on LLM judges or human panels -- both subjective and disconnected from actual research impact. We introduce HindSight, a time-split evaluation framework that measures idea quality by matching generated ideas against real future publications and scoring them by citation impact and venue acceptance. Using a temporal cutoff~$T$, we restrict an idea generation system to pre-$T$ literature, then evaluate its outputs against papers published in the subsequent 30 months. Experiments across 10 AI/ML research topics reveal a striking disconnect: LLM-as-Judge finds no significant difference between retrieval-augmented and vanilla idea generation ($p{=}0.584$), while HindSight shows the retrieval-augmented system produces 2.5$\times$ higher-scoring ideas ($p{<}0.001$). Moreover, HindSight scores are \emph{negatively} correlated with LLM-judged novelty ($ρ{=}{-}0.29$, $p{<}0.01$), suggesting that LLMs systematically overvalue novel-sounding ideas that never materialize in real research.

研究创新大模型评估真实影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。