用强化学习任务评估并提升大模型在药物设计中的能力
Evaluating the Progression of Large Language Model Capabilities for Small-Molecule Drug Design
- 构建化学任务作为强化学习环境,统一评估与微调
- 小样本实验中模型表现仍有较大提升空间
- 微调后小模型可媲美顶尖大模型,适合资源有限团队
大语言模型(LLMs)具备从多元信息源中推理的能力,有望加速小分子药物设计。然而其实际效用尚不明确,因缺乏反映真实场景的评测基准。本文提出一套涵盖分子性质预测、分子表征转换与分子设计的化学相关任务,并将这些任务建模为强化学习(RL)环境,实现评估与后训练的统一。在三个模型家族上测试发现,前沿模型在化学任务上日益熟练,但在低数据实验条件下仍有显著提升空间。关键的是,基于RL的后训练可大幅改善性能:一个较小模型经本环境微调后,表现可与最先进前沿模型相当,尽管其基础模型能力较弱。这表明通过精心设计的评测任务与针对性后训练,可有效揭示并弥补关键能力缺口,为药物发现中实用化LLM提供可行路径。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have the potential to accelerate small molecule drug design due to their ability to reason about information from diverse sources and formats. However, their practical utility remains unclear due to the lack of benchmarks that reflect real-world scenarios. In this work, we introduce a suite of chemically-grounded tasks spanning molecular property prediction, molecular representation transformations, and molecular design. Importantly, we formulate these tasks as reinforcement learning (RL) environments, enabling a unified approach for evaluation and post-training. Across three model families, we find that frontier models are increasingly proficient at chemical tasks, but that there is significant room for improvement, especially in experimental settings with low data. Critically, we show that RL-based post-training can substantially improve performance. A smaller model post-trained on our environments becomes competitive with state-of-the-art frontier models, despite a significantly weaker base model. This suggests a practical route toward employing LLMs in drug discovery; by combining carefully-designed evaluation tasks with targeted post-training, we can both elucidate and close critical capability gaps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。