用强化学习训练出可自然语言推理的化学模型,效果超人类专家。
Training a Scientific Reasoning Model for Chemistry
- 基于64万道实验数据问题,用强化学习微调240亿参数模型
- 在375个化学任务中超越通用模型与人类专家,数据效率更高
- 适合需要高效建模的科研人员,推动跨科学领域应用
推理模型是通过生成长链思维过程来回答问题的大语言模型,兼具高准确率和可解释性。现有研究主要集中在数学、编程和逻辑领域,但其是否适用于化学等自然科学尚不明确。本文展示无需额外领域预训练即可对语言模型进行化学推理能力的后训练,且所需数据量远低于当前专用模型。我们提出ether0,一个基于Mistral-Small-24B的240亿参数大模型,能以自然语言进行推理并输出化学结构。该模型在640,730道实验基础的化学问题上,通过强化学习训练,覆盖从可合成性、血脑屏障渗透性、人体受体活性到气味等共375项任务。结果显示,ether0在分子设计任务上优于通用化学模型、前沿模型及人类专家,且数据效率显著提升。本方法有望推广至其他科学领域的高效语言模型构建。
原文摘要 · Abstract (English)
Reasoning models are large language models that emit a long chain-of-thought before answering, providing both higher accuracy and explicit reasoning for their response. A major question has been whether language model reasoning generalizes beyond mathematics, programming, and logic, where most previous work has focused. We demonstrate that reasoning models can be post-trained for chemistry without additional domain pretraining, and require substantially less data compared to contemporary domain-specific models. We report ether0, a 24B parameter LLM (based on Mistral-Small-24B) that can reason in natural language and respond with chemical structures. This reasoning model was trained with reinforcement learning on 640,730 experimentally-grounded chemistry problems across 375 tasks ranging from synthesizability, to blood-brain barrier permeability, to human receptor activity, to scent. Our model exceeds general-purpose chemistry models, frontier models, and human experts on molecular design tasks. It is also more data efficient relative to specialized models. We anticipate that this method can be applied to train data-efficient language models specialized for tasks across a wide variety of scientific domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。