arXiv:2608.22948cs.CLcs.AI2026-08

用可证伪实验设计评估大模型科研创意,让判断有标准可依。

What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation

论文配图:What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation
图 1 · 摘自论文原文
  • 提出六领域契约式评测框架,要求每条创意预先定义反例。
  • 四模型在1200次盲评中形成稳定排名,优劣由测试设计质量决定。
  • 适用于需严谨评估生成创意的科研与AI研发人员。

大语言模型被越来越多用于生成科研想法,但现有评判方式缺乏统一标准:自由形式评价受风格和立场影响,按后续论文评分则仅奖励已实现路径的重现。本文提出一个从文献到验证的基准测试(Lit2Test),基于六领域契约结构,要求每个研究设想预先承诺能证明其错误的观测结果,使创意质量可判定而非仅可争论。该基准前瞻性构建于200个真实论文邻域,对四个前沿模型生成的创意进行1200次双盲对比评估。通过诊断控制与有限人类校准,三名标注者在明确可靠性范围内达成一致结论。在全部10,000次自助重采样中,四模型均保持严格排序,差异源于所提测试与度量的质量,而非表面流畅性。本文开源基准、构建流程及审计材料。

原文摘要 · Abstract (English)

Large language models are increasingly used to propose research ideas, yet the prevailing ways of judging such ideas supply no shared decision rule: free-form judging sways with style and position, and scoring against a later paper rewards recovery of one realized trajectory. We introduce a benchmark that carries a proposal from Literature to Test: the Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable. Built prospectively from 200 real-paper neighborhoods, Lit2Test elicits proposals from four frontier models and compares them through 1,200 pairwise comparisons judged blind in both presentation orders. The protocol audits its own reliability through diagnostic controls and bounded human calibration, with three annotators corroborating the conclusions within explicitly stated reliability bounds. Lit2Test recovers a strict ranking of the four models in all 10,000 bootstrap replicates, and the separation comes from the quality of the proposed tests and metrics rather than from surface fluency. We release the benchmark, construction pipeline, and audit artifacts for public use.

科研创新模型评测可证伪性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。