arXiv:2604.13201cs.CLcs.AI2026-04中稿 · COLM

构建可无限生成的科学数据基准,评估模型真实科研推理能力

InfiniteScienceGym: An Unbounded, Procedurally-Generated Benchmark for Scientific Analysis

论文配图:InfiniteScienceGym: An Unbounded, Procedurally-Generated Benchmark for Scientific Analysis
图 1 · 摘自论文原文
  • 用程序化生成技术创建带真实数据结构的科学仓库和可验证问答任务
  • 模型整体准确率不足50%,识别无法回答问题仍是主要短板
  • 适合评估模型在真实科研场景中的推理、拒答与工具使用能力

大型语言模型正成为科学助手,但评估其基于实证数据的推理能力仍具挑战。现有基于已发表研究和人工标注的基准存在发表偏见、已知知识偏见、标签噪声及存储开销大等问题。我们提出InfiniteScienceGym,一个通过种子程序化生成的科学数据仓库基准,配套可验证的问答任务。该模拟器可确定性生成包含真实目录结构、文件与表格数据的自包含仓库,并由特权问答生成器产生有确切答案和无解问题。这使得在不依赖大规模静态数据集的前提下,可控地评估证据基础推理、拒答行为与工具辅助分析。该基准补充了真实科学数据集的盲点与失效模式。对专有与开源模型的评估显示,所有模型整体准确率均未超过50%,识别无法回答问题仍是主要弱点,且更强模型更善于使用工具而非单纯增加输入长度。

原文摘要 · Abstract (English)

Large language models are emerging as scientific assistants, but evaluating their ability to reason from empirical data remains challenging. Benchmarks derived from published studies and human annotations inherit publication bias, known-knowledge bias, label noise, and substantial storage requirements. We present InfiniteScienceGym, a procedurally generated benchmark of scientific repositories paired with a verifiable question-answering task. From a seed, the simulator deterministically generates a self-contained repository with realistic directory structure, files, and tabular data, and a privileged QA generator produces both answerable and unanswerable questions with exact ground truth. This makes it possible to evaluate evidence-grounded reasoning, abstention, and tool-mediated analysis in a controlled setting without distributing a large static corpus. InfiniteScienceGym complements real scientific benchmarks by targeting blind spots and failure modes that are hard to evaluate using published datasets alone. Evaluating both proprietary and open-weight models, we find that none achieve more than 50% accuracy overall, that recognizing unanswerable questions remains a major weakness, and that stronger models tend to use tools more effectively rather than simply consuming more tokens.

科学推理数据生成基准测试语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。