用自然语言指挥多智能体系统,一键生成并筛选新药候选分子
MADD: Multi-Agent Drug Discovery Orchestra
- 四个协作智能体分工处理从零生成到筛选的全流程
- 在7个靶点上表现优于现有AI方案,成功发现5个新靶点的候选分子
- 适合非编程背景的实验研究人员快速开展AI驱动的新药设计
先导化合物识别是早期药物研发的核心挑战,传统方法依赖大量实验资源。近年来,大语言模型(LLMs)推动了虚拟筛选技术的发展,降低了成本并提升了效率。然而,工具复杂度上升限制了湿实验研究人员的使用。多智能体系统通过结合LLM的可解释性与专用模型的精度,提供可行解决方案。本文提出MADD,一个能根据自然语言查询构建并执行定制化先导化合物识别流程的多智能体系统。MADD包含四个协同智能体,分别负责从头化合物生成与筛选的关键任务。我们在七个药物发现案例中评估MADD,结果表明其性能优于现有基于LLM的方案。利用MADD,我们首次实现对五个生物靶点的AI原生设计,并公开了发现的先导分子。此外,我们构建了一个包含超过三百万种化合物的查询-分子对及对接评分的新基准数据集,以推动药物设计的智能体化发展。
原文摘要 · Abstract (English)
Hit identification is a central challenge in early drug discovery, traditionally requiring substantial experimental resources. Recent advances in artificial intelligence, particularly large language models (LLMs), have enabled virtual screening methods that reduce costs and improve efficiency. However, the growing complexity of these tools has limited their accessibility to wet-lab researchers. Multi-agent systems offer a promising solution by combining the interpretability of LLMs with the precision of specialized models and tools. In this work, we present MADD, a multi-agent system that builds and executes customized hit identification pipelines from natural language queries. MADD employs four coordinated agents to handle key subtasks in de novo compound generation and screening. We evaluate MADD across seven drug discovery cases and demonstrate its superior performance compared to existing LLM-based solutions. Using MADD, we pioneer the application of AI-first drug design to five biological targets and release the identified hit molecules. Finally, we introduce a new benchmark of query-molecule pairs and docking scores for over three million compounds to contribute to the agentic future of drug design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。