用历史回测法评估科学问题生成系统,让时间检验其前瞻性。
Historical Backtesting for Scientific Question Discovery: A Protocol and Astronomy Pilot
- 构建时间隔离的回测协议,冻结问题与文献,由未来数据判断问题价值。
- 无模型生成器在所有时期均发现被未来证伪的假设,优于仅靠大模型提示的生成。
- 多模型一致度高但人类标注者分歧大,证明评估体系需依赖模型而非人工。
当前科学问题生成系统的评估依赖专家评分、大模型评判或精选案例,均为主观且不可证伪。本文提出历史回测新范式:将语料库冻结于历史时点,系统生成问题并冻结,再通过时间隔离的未来语料库判断每个问题是否被回答、部分解决、独立提出或被忽略,以及其前提是否被支持或反驳。该协议与模型无关,任何生成冻结问题的系统均可评估。研究发布可复现的天文学实例,包含时间隔离语料、冻结问题、可审计标签、四类基准方法和提交接口。两个核心发现:第一,基于证据结构的生成方法优于仅使用大模型提示的方法;在2010-2024年四个时间窗口(共798个问题)的交叉测试中,后者表现出训练数据的记忆性,缺乏前瞻性;而完全不使用模型权重的生成器,在所有时期均发现被未来证伪的前提。第二,七人评审研究(两名盲评人类、五名模型判官,90个样本)显示,评估分类体系本身存在歧义,而非判官问题:两位人类标注者一致性仅为kappa=0.17,各模型判官与专业标注者一致率在0.17-0.26之间,而前沿模型间一致性达0.60;若仅以模型间一致性为标准,将高估模型可靠性三倍。研究还发布了前瞻性实例:200个问题于2026年8月17日冻结,将在2027-2030年期间被评分,使核心结论接受无污染的时间检验。
原文摘要 · Abstract (English)
Systems that generate scientific research questions are evaluated today by expert scores, LLM-as-judge ratings, or curated case studies -- all subjective, none falsifiable. We formalize historical backtesting as an alternative: a system generates questions from a corpus frozen at a historical cutoff, the questions are frozen before any access to later literature, and a temporally isolated future corpus then determines whether each question was subsequently answered, partially addressed, independently posed, or ignored, and whether its underlying premise was supported or refuted. The protocol is model-agnostic: any system that emits frozen questions can be scored. We release reproducible astronomy instances with temporally isolated corpora, frozen questions, auditable labels, four reference baselines, and a submission interface. Two findings result. First, evidence-structure-first generation outperforms LLM-only prompting: across a generator decomposition crossed with a four-cutoff stress test (2010-2024, 798 judged questions) whose last window postdates model training, LLM-only generation shows memorized relevance without specific foresight, while a generator using no model weights at all finds questions whose premises the future refutes in every era. Second, a seven-rater agreement study (two blinded human annotators, five judge models, 90 items) indicts the outcome taxonomy rather than the judge: two careful humans agree at kappa = 0.17, every judge model agrees with the professional annotator as well or better (0.17-0.26), and frontier models agree with one another at 0.60 -- certifying an LLM judge by model-model agreement would have overstated its reliability threefold. A prospective instance -- 200 questions frozen 2026-08-17, scored 2027-2030 -- is released so the central claims become contamination-free tests that time itself will grade.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。