arXiv:2608.16645cs.AIcs.CL2026-08

测试大模型能否从论文参考文献中还原研究核心思想。

Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies

  • 用盲测设计隔离论文正文,仅凭参考文献推断研究主题。
  • 七款前沿模型平均准确率仅3%-15%,表明还原难度极高。
  • 多模型协作+锦标赛机制提升至23%-42%,适合研究方法论探索者。

我们提出Reconstruction,一个盲测型研究思想还原基准,仅提供论文的预发表参考文献,不包含原始论文及同期或未来文献。通过严格防泄露协议——时间引用截止、匿名参考文献编号、冻结每篇论文的参考文献列表,防止在提示时泄露原始研究思想。在六个科学领域共643篇论文上评估,七个前沿语言模型的匹配率仅为约3%-15%。随后我们测试了一种仅依赖参考文献的多智能体(前四名)流水线,结合跨模型评审与对齐假设槽的瑞士轮淘汰赛制,无需外部网络搜索。该方法使匹配率提升至约23%-42%,相较最优单模型基线提高约2.4倍。本稿报告了实验协议、防泄露设计及当前结果,作为arXiv时间戳发布。

原文摘要 · Abstract (English)

Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We introduce Reconstruction, a blind idea-recovery benchmark that withholds the seed paper and all contemporaneous or future literature, and asks models to propose hypotheses that an independent large language model judge matches against the held-out ground-truth idea. A strict anti-leakage protocol-temporal citation cutoff, anonymous reference IDs, and frozen per-paper bibliographies, which prevents prompt-time leakage of the seed idea. Across six scientific domains and 643 evaluated papers, seven frontier models achieve only modest Match rates (approx. 3-15%). We then evaluate a reference-only multi-agent (top 4) pipeline that combines cross-model review with a Swiss tournament over aligned hypothesis slots, without external web search. Cross-model review plus tournament selection raises Match rates to approx. 23-42% across all six domains, which is an observed approx. 2.4x lift over the best single-model baseline. This draft reports the protocol, anti-leakage design, and current results as an arXiv timestamp.

语言模型盲测知识还原推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。