arXiv:2410.13147cs.LGcs.AI2024-10EMNLP被引 4

用智能体流程让大模型更准地修改分子结构,提升药物研发效率。

AgentDrug: Utilizing Large Language Models in An Agentic Workflow for Zero-Shot Molecular Editing

  • 构建嵌套循环的智能体工作流,结合化学工具反馈优化分子结构。
  • 在单/多属性编辑任务中,准确率提升最高达29%,优于现有方法。
  • 适用于零样本分子改造,适合药物研发与生成式化学研究者。

分子编辑——通过修改现有分子以提升期望性质——是药物发现中的核心任务。尽管大语言模型(LLM)可借助自然语言实现编辑,但直接提示法准确率有限。本文提出AgentDrug,一种基于智能体的工作流,通过结构化迭代优化显著提升准确性。该流程包含嵌套循环:内层利用化学信息学工具验证分子结构,外层则通过通用反馈与基于梯度的目标引导模型向性质优化方向演进。我们在单属性和多属性编辑基准上进行评估,设定宽松与严格阈值。结果表明,使用Qwen-2.5-3B模型时,六项单属性任务准确率分别提升20.7%(宽松)和16.8%(严格),八项多属性任务提升7.0%和5.3%;使用更大模型Qwen-2.5-7B时,单属性任务提升达28.9%(宽松)和29.0%(严格),多属性任务提升14.9%(宽松)和13.2%(严格)。

原文摘要 · Abstract (English)

Molecular editing-modifying a given molecule to improve desired properties-is a fundamental task in drug discovery. While LLMs hold the potential to solve this task using natural language to drive the editing, straightforward prompting achieves limited accuracy. In this work, we propose AgentDrug, an agentic workflow that leverages LLMs in a structured refinement process to achieve significantly higher accuracy. AgentDrug defines a nested refinement loop: the inner loop uses feedback from cheminformatics toolkits to validate molecular structures, while the outer loop guides the LLM with generic feedback and a gradient-based objective to steer the molecule toward property improvement. We evaluate AgentDrug on benchmarks with both single- and multi-property editing under loose and strict thresholds. Results demonstrate significant performance gains over previous methods. With Qwen-2.5-3B, AgentDrug improves accuracy by 20.7% (loose) and 16.8% (strict) on six single-property tasks, and by 7.0% and 5.3% on eight multi-property tasks. With larger model Qwen-2.5-7B, AgentDrug further improves accuracy on 6 single-property objectives by 28.9% (loose) and 29.0% (strict), and on 8 multi-property objectives by 14.9% (loose) and 13.2% (strict).

分子编辑智能体大模型药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。