用可验证数据训练生物领域专用大模型,性能超越商用大模型。
OwkinZero: Accelerating Biological Discovery with AI
- 通过可验证问答对微调,用强化学习提升模型生物推理能力。
- 8-32B规模模型在药物靶点、作用机制等任务上优于更大商用模型。
- 专精模型在未见任务上表现更优,混合数据训练增强跨任务泛化。
尽管大语言模型(LLMs)快速推动科学研究,但在药物发现中的靶点可成药性、药物作用方式适配性、药物扰动效应等核心生物推理任务上仍表现不足。为此,我们构建并整理了八个综合性基准数据集,包含超过30万条可验证的问答对,覆盖关键药物发现挑战。基于此资源,我们通过可验证奖励的强化学习策略,对开源大模型进行后训练,开发出OwkinZero系列模型。结果表明,8-32B规模的OwkinZero模型在这些生物基准测试中显著优于更大的先进商业大模型。值得注意的是,我们在单任务上训练的专精模型在未见过的任务上也持续优于其基础模型,展现出关键泛化特性;而综合多数据集训练的OwkinZero模型进一步放大该效果,实现更广泛的跨任务提升。本研究标志着解决当前大模型生物推理盲区的重要进展,证明在精心策划数据上进行定向强化学习,可释放专精模型的可泛化性能,加速人工智能驱动的生物学发现。
原文摘要 · Abstract (English)
While large language models (LLMs) are rapidly advancing scientific research, they continue to struggle with core biological reasoning tasks essential for translational and biomedical discovery. To address this limitation, we created and curated eight comprehensive benchmark datasets comprising over 300,000 verifiable question-and-answer pairs, each targeting critical challenges in drug discovery including target druggability, modality suitability, and drug perturbation effects. Using this resource, we developed the OwkinZero models by post-training open-source LLMs through a Reinforcement Learning from Verifiable Rewards strategy. Our results demonstrate that specialized 8-32B OwkinZero models substantially outperform larger, state-of-the-art commercial LLMs on these biological benchmarks. Remarkably, we uncover evidence of a key aspect of generalization: specialist models trained on a single task consistently outperform their base models on previously unseen tasks. This generalization effect is further amplified in our comprehensive OwkinZero models, which were trained on a mixture of datasets and achieve even broader cross-task improvements. This study represents a significant step toward addressing the biological reasoning blind spot in current LLMs, demonstrating that targeted reinforcement learning on carefully curated data can unlock generalizable performance in specialized models, thereby accelerating AI-driven biological discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。