arXiv:2602.20459cs.AIcs.CL2026-02被引 1

构建首个科学预测数据集,评估AI能否预判未来科研进展。

PreScience: A Dataset and Benchmark for Scientific Forecasting

  • 基于9.8万篇AI论文构建包含50.2万篇文献的时序数据集
  • 7个任务验证模型对论文贡献、引用、合作等的预测能力
  • 发现AI生成论文创新性与多样性均低于人类研究

现有科学记录能否训练出能预测未来科研进展的AI系统?我们提出PreScience,一个围绕98,000篇近期人工智能研究论文构建的数据集与基准测试,辅以作者发表历史和引文链接,共涵盖502,000篇论文。数据记录包括标题、摘要、去重作者身份、重要参考文献、主题标签、引文轨迹及时间截断元数据。我们设计了七个典型任务:五个以论文为中心的任务——贡献生成、合作者预测、前期工作选择、引用次数预测、未来组合预测,以及两个聚合主题趋势预测变体。我们开发了从简单启发式到前沿语言模型和代理系统的基线方法,并引入LACER——一种基于LLM的生成贡献描述相似性评估指标,其与人工判断一致性优于现有指标。最后,我们用任务模型生成12个月合成论文语料,发现生成论文在创新性和多样性上系统性低于同期人类研究。数据集(https://huggingface.co/datasets/allenai/prescience)和代码(https://github.com/allenai/prescience)已开源。

原文摘要 · Abstract (English)

Can AI systems trained on the existing scientific record forecast the advances that will follow? We introduce PreScience, a dataset and benchmark for scientific forecasting built around 98K recent AI research papers, together with companion papers covering author publication histories and citation links, yielding 502K papers in total. The resulting paper records include titles, abstracts, disambiguated author identities, influential references, topic labels, citation trajectories, and metadata snapshotted to respect temporal cutoffs. We instantiate seven exemplar tasks: five paper-anchored tasks -- contribution generation, collaborator prediction, prior work selection, citation count prediction, and future combination prediction -- and two aggregate topic trend forecasting variants. We develop baselines ranging from simple heuristics and embedding methods to frontier language models and agentic systems, and introduce LACER, an LLM-based metric for evaluating similarity of generated contribution descriptions that agrees better with human judgments than existing metrics. Finally, we compose task models to generate a 12-month synthetic corpus and find that the resulting papers are systematically less diverse and less novel than human-authored research from the same period. We release the PreScience dataset (https://huggingface.co/datasets/allenai/prescience) and code (https://github.com/allenai/prescience).

科学预测数据集LLM评估AI科研

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。