arXiv:2508.03444cs.LG2025-08被引 4

用多智能体系统让药物分子设计可追溯、可验证,提升优化效率。

An Auditable Agent Platform For Automated Molecular Optimisation

  • 构建分层智能体框架,分工协作完成从设计到评估的全流程。
  • 多智能体模式使结合亲和力平均提升31%,优于单智能体和纯大模型方案。
  • 每步操作留痕,实现全过程可审计,适合需要透明性的药物研发团队。

药物发现常因数据、知识与工具分散而停滞。为此,我们构建了一个分层智能体框架,自动化分子优化流程:主研究员设定目标,数据库智能体获取靶点信息,AI专家使用序列到分子的深度学习模型生成新骨架,药物化学家编辑并调用对接工具,评分智能体打分,科学批评家审查逻辑。每次工具调用均记录摘要,确保推理路径全程可追溯。智能体通过简洁的溯源记录通信,追踪分子演化轨迹,并通过上下文学习复用成功转化。在AKT1蛋白上,基于五个大语言模型运行了三轮研究循环。按平均对接得分排序后,对前两名模型进行了20次独立扩展测试。比较三种配置下顶尖LLM的表现:仅用LLM、单智能体、多智能体。结果表明,多智能体架构在聚焦结合优化方面表现最优,平均预测结合亲和力提升31%;单智能体生成分子药学性质更优,但结合得分较低;无引导的LLM运行最快,但缺乏透明的工具信号,其推理路径有效性无法验证。这些结果说明,推理时扩展、聚焦反馈环与溯源机制可将通用大模型转化为可审计的分子设计系统,未来拓展至ADMET与选择性预测工具,将进一步推进发现流程。

原文摘要 · Abstract (English)

Drug discovery frequently loses momentum when data, expertise, and tools are scattered, slowing design cycles. To shorten this loop we built a hierarchical, tool using agent framework that automates molecular optimisation. A Principal Researcher defines each objective, a Database agent retrieves target information, an AI Expert generates de novo scaffolds with a sequence to molecule deep learning model, a Medicinal Chemist edits them while invoking a docking tool, a Ranking agent scores the candidates, and a Scientific Critic polices the logic. Each tool call is summarised and stored causing the full reasoning path to remain inspectable. The agents communicate through concise provenance records that capture molecular lineage, to build auditable, molecule centered reasoning trajectories and reuse successful transformations via in context learning. Three cycle research loops were run against AKT1 protein using five large language models. After ranking the models by mean docking score, we ran 20 independent scale ups on the two top performers. We then compared the leading LLMs' binding affinity results across three configurations, LLM only, single agent, and multi agent. Our results reveal an architectural trade off, the multi agent setting excelled at focused binding optimization, improving average predicted binding affinity by 31%. In contrast, single agent runs generated molecules with superior drug like properties at the cost of less potent binding scores. Unguided LLM runs finished fastest, yet their lack of transparent tool signals left the validity of their reasoning paths unverified. These results show that test time scaling, focused feedback loops and provenance convert general purpose LLMs into auditable systems for molecular design, and suggest that extending the toolset to ADMET and selectivity predictors could push research workflows further along the discovery pipeline.

药物设计智能体系统可审计性分子优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。