arXiv:2602.08213cs.LGcs.AI2026-02被引 1

用大模型显式推理优化药物分子,提升药效同时保持结构相似性。

DrugR: Optimizing Molecular Drugs through LLM-based Explicit Reasoning

  • 通过分步药理推理引导药物分子优化
  • 显著提升ADMET多维度性能且不降低结合亲和力
  • 结果可解释,适合药物研发人员参考

分子生成与优化是化学领域的基础任务。尽管大语言模型(LLMs)具备丰富知识和交互能力,但其在分子结构与药理性质间复杂隐含关系的理解上仍受限,且缺乏标注数据。为此,我们提出DrugR,一种将显式、分步药理推理引入优化过程的LLM方法。该方法融合领域特定的持续预训练、基于逆向数据工程的监督微调,以及自平衡多粒度强化学习。实验表明,DrugR在不损害原始分子核心活性和结构相似性的前提下,全面提升了多项ADMET属性。其显式推理过程可提供清晰可解释的优化依据,生成可操作的设计洞见,推动自动化、知识驱动的科学发现。代码与模型检查点已开源。

原文摘要 · Abstract (English)

Molecule generation and optimization is a fundamental task in chemical domain. The rapid development of intelligent tools, especially large language models (LLMs) with powerful knowledge reserves and interactive capabilities, has provided new paradigms for it. Nevertheless, the intrinsic challenge for LLMs lies in the complex implicit relationship between molecular structure and pharmacological properties and the lack of corresponding labeled data. To bridge this gap, we propose DrugR, an LLM-based method that introduces explicit, step-by-step pharmacological reasoning into the optimization process. Our approach integrates domain-specific continual pretraining, supervised fine-tuning via reverse data engineering, and self-balanced multi-granular reinforcement learning. This framework enables DrugR to effectively improve key ADMET properties while preserving the original molecule's core efficacy. Experimental results demonstrate that DrugR achieves comprehensive enhancement across multiple properties without compromising structural similarity or target binding affinity. Importantly, its explicit reasoning process provides clear, interpretable rationales for each optimization step, yielding actionable design insights and advancing toward automated, knowledge-driven scientific discovery. Our code and model checkpoints are open-sourced to foster future research.

药物生成大模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。