arXiv:2601.05567cs.AIcs.CL2026-01ACL被引 3

构建跨9个学科的科学推理数据集,推动大模型在真实文献中的科学问答能力。

WildSci: Advancing Scientific Reasoning from In-the-Wild Literature

  • 从真实文献自动构建多领域科学选择题,支持可扩展训练。
  • 用强化学习微调后,在多个科学基准上显著提升推理性能。
  • 适合研究科学推理、AI for Science 的学者和开发者使用。

大型语言模型(LLM)在数学和编程等领域的推理进展迅速,得益于高质量数据和客观评估指标。然而,在医学、材料科学等科学领域,由于数据覆盖有限且问题开放性强,进展仍受制约。为此,我们提出WildSci,一个从同行评审文献中自动生成的领域特定科学问题数据集,涵盖9个科学领域和26个子领域。通过将复杂科学推理任务转化为多选题形式,实现可扩展训练与明确奖励信号。我们进一步采用强化学习对模型进行微调,并分析训练动态,包括各领域表现变化、回答行为及泛化趋势。在一系列科学基准上的实验验证了该数据集与方法的有效性。我们已开源WildSci,以促进科学推理的可持续研究,地址为https://huggingface.co/datasets/JustinTX/WildSci。

原文摘要 · Abstract (English)

Recent progress in large language model (LLM) reasoning has focused on domains like mathematics and coding, where abundant high-quality data and objective evaluation metrics are readily available. In contrast, progress in LLM reasoning models remains limited in scientific domains such as medicine and materials science due to limited dataset coverage and the inherent complexity of open-ended scientific questions. To address these challenges, we introduce WildSci, a new dataset of domain-specific science questions automatically synthesized from peer-reviewed literature, covering 9 scientific disciplines and 26 subdomains. By framing complex scientific reasoning tasks in a multiple-choice format, we enable scalable training with well-defined reward signals. We further apply reinforcement learning to finetune models on these data and analyze the resulting training dynamics, including domain-specific performance changes, response behaviors, and generalization trends. Experiments on a suite of scientific benchmarks demonstrate the effectiveness of our dataset and approach. We release WildSci to enable scalable and sustainable research in scientific reasoning, available at https://huggingface.co/datasets/JustinTX/WildSci.

科学推理大模型数据集强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。