用未来可验证的科研成果评估大模型判断科学想法的能力。
Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments
- 构建时间分段的离线沙盒,让模型预测未来可验证的科研影响。
- 在3万多个案例中发现,工具使用效果因任务而异,交互预算越高表现越好。
- 适合研究智能体评估科学创意、人机判断差异的学者使用。
大型语言模型在评估和预测科研想法方面日益重要,但缺乏可扩展的方法来衡量其判断质量。为此,我们提出PoT,一个半可验证的基准框架,将科学想法的判断与后续可观察的下游信号(如引用量、研究议程变化)关联起来。PoT冻结某一截止时间前的证据快照,在离线沙盒中要求模型预测截止时间后的结果,待真实情况出现后即可验证,实现无需大量专家标注的可扩展评估,并能分析人类与模型在同行评审奖项等信号上的判断偏差。此外,PoT还提供了一个可控的测试环境,用于评估基于智能体的科研判断,比较使用工具的智能体与不使用工具的基线在提示扰动和预算扩展下的表现。在涵盖四个领域、超过3万实例的数据上,我们发现:相比非智能体基线,更高的交互预算通常提升智能体表现,而工具使用的收益则高度依赖具体任务。通过结合时间划分的未来可验证目标与离线工具使用沙盒,PoT支持对面向未来的科学想法判断任务中智能体的可扩展评估。
原文摘要 · Abstract (English)
Large language models are increasingly being used to assess and forecast research ideas, yet we lack scalable ways to evaluate the quality of models' judgments about these scientific ideas. Towards this goal, we introduce PoT, a semi-verifiable benchmarking framework that links scientific idea judgments to downstream signals that become observable later (e.g., citations and shifts in researchers' agendas). PoT freezes a pre-cutoff snapshot of evidence in an offline sandbox and asks models to forecast post-cutoff outcomes, enabling verifiable evaluation when ground truth arrives, scalable benchmarking without exhaustive expert annotation, and analysis of human-model misalignment against signals such as peer-review awards. In addition, PoT provides a controlled testbed for agent-based research judgments that evaluate scientific ideas, comparing tool-using agents to non-agent baselines under prompt ablations and budget scaling. Across 30,000+ instances spanning four benchmark domains, we find that, compared with non-agent baselines, higher interaction budgets generally improve agent performance, while the benefit of tool use is strongly task-dependent. By combining time-partitioned, future-verifiable targets with an offline sandbox for tool use, PoT supports scalable evaluation of agents on future-facing scientific idea judgment tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。