arXiv:2604.03506cs.AI2026-04

用生物文献构建可验证的推理数据,提升模型在生物学问题上的推理能力。

BioAlchemy: Distilling Biological Literature into Reasoning-Ready Reinforcement Learning Training Data

  • 从生物科研文本中自动提取可验证的问答对,构建多样化数据集。
  • 新数据集使模型在生物推理任务上提升9.12%准确率。
  • 适合想提升生物领域推理能力的研究者和开发者。

尽管生物学训练文本体量庞大,但推理模型在该领域的应用仍落后于数学和编程。本文发现,现有大规模推理数据集中的生物问题与现代生物学研究主题分布不匹配,这种主题失衡可能影响模型性能。同时,我们指出从生物科研文本中提取具有挑战性且可验证的研究问题,是利用强化学习提升生物学任务表现的关键但未被充分开发的环节。为此,我们提出 BioAlchemy 流程,从生物科研语料中抽取多样化的可验证问答对,构建包含超过 345,000 个科学推理问题的 BioAlchemy-345K 数据集。进一步证明,将数据集与现代生物学主题分布对齐后,结合强化学习可显著提升推理性能。最后,我们推出 BioAlchemist-8B 模型,在生物基准测试中相比基础模型提升 9.12%。结果表明该方法能有效增强模型在生物学中的科学推理能力。模型已开源:https://huggingface.co/BioAlchemy。

原文摘要 · Abstract (English)

Despite the large corpus of biology training text, the impact of reasoning models on biological research generally lags behind math and coding. In this work, we show that biology questions from current large-scale reasoning datasets do not align well with modern research topic distributions in biology, and that this topic imbalance may negatively affect performance. In addition, we find that methods for extracting challenging and verifiable research problems from biology research text are a critical yet underdeveloped ingredient in applying reinforcement learning for better performance on biology research tasks. We introduce BioAlchemy, a pipeline for sourcing a diverse set of verifiable question-and-answer pairs from a scientific corpus of biology research text. We curate BioAlchemy-345K, a training dataset containing over 345K scientific reasoning problems in biology. Then, we demonstrate how aligning our dataset to the topic distribution of modern scientific biology can be used with reinforcement learning to improve reasoning performance. Finally, we present BioAlchemist-8B, which improves over its base reasoning model by 9.12% on biology benchmarks. These results demonstrate the efficacy of our approach for developing stronger scientific reasoning capabilities in biology. The BioAlchemist-8B model is available at: https://huggingface.co/BioAlchemy.

生物推理强化学习数据构建科学智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。