arXiv:2606.13020cs.AI2026-06

构建可调控的科学推理基准,评估大模型在三类推理上的表现。

SciR: A Controllable Benchmark for Scientific Reasoning in LLMs

  • 基于形式化对象生成任务,保证答案可验证
  • 独立控制信息提取与推理难度,揭示模型瓶颈
  • 适合研究科学推理能力或模型弱点的学者使用

科学推理中反复出现三种典型推理范式:演绎、归纳和因果溯因。当前可靠评估大模型在科学场景下的这些推理能力仍面临挑战:人工标注的科学基准成本高且缺乏机制性真值,而合成逻辑推理基准又不贴近真实科学文献。我们提出SciR,一个融合多范式推理与可控科学文本生成的基准,基于三个典型的科学问题构建。任务从形式化对象(演绎树、归纳规则假设、因果图)生成,确保答案可验证,再通过各领域定制的文体渲染为多文档科学论述。该设计允许独立调节两个难度维度:信息提取的难易程度,以及推理本身的复杂度。我们测试了六种模型,发现两个维度均对所有模型造成影响,且效果叠加。令人意外的是,连将推理交给已验证求解器的神经符号系统也受文本渲染负面影响。每种模型呈现出独特的‘提取-推理’表现分布,例如DeepSeek-R1等推理模型在推理维度上显著优于非推理指令模型。据我们所知,SciR是首个在提取与推理难度上均可参数化控制的多范式科学推理基准。

原文摘要 · Abstract (English)

Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction. Reliably evaluating LLMs on these in scientific settings is currently out of reach: scientific benchmarks built on human annotations are costly and lack mechanistic ground truth, while synthetic logical-reasoning benchmarks do not resemble real scientific documents. We introduce SciR, a benchmark that combines multi-paradigm reasoning with controllable scientific rendering, anchored on three paradigmatic scientific problems. Tasks are generated from formal objects (deduction tree, inductive rule hypothesis, causal graph) to guarantee verifiable answers, then rendered into multi-document scientific discourse via per-track domain-tuned genres. The construction lets us independently vary two difficulty axes: how hard it is to extract the key information needed for inference, and how hard the principled inference itself is. We test six models. Both axes hurt every model, and their effects compound. The rendering even hurts neurosymbolic pipelines, which hand inference to a verified solver. The two axes yield a per-model extraction-vs-inference profile: for instance, reasoning models like deepseek-r1 mostly surpass non-reasoning instruct models on the inference axis. To our knowledge, SciR is the first multi-paradigm scientific-reasoning benchmark with parametric control on both extraction and inference difficulty.

科学推理大模型评测可控生成多范式推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。