arXiv:2505.13940cs.AIq-bio.BM2025-05被引 24

DrugPilot用参数化记忆实现药物发现全流程自动化

DrugPilot: LLM-based Parameterized Reasoning Agent for Drug Discovery

  • 设计参数化记忆池统一处理多源异构数据
  • 在多工具、多轮对话场景下任务完成率超90%
  • 适合需要交互式科学推理的计算药物研发人员

大型语言模型(LLMs)与自主代理结合在科学发现中展现出巨大潜力,但应用于药物发现仍受限于大规模多模态数据处理、任务自动化不足及领域工具支持弱等问题。为此,我们提出DrugPilot,一种基于LLM的参数化推理代理系统,专为药物发现端到端科研流程设计。该系统通过结构化工具调用与新型参数化记忆池集成,将公开数据与用户输入的异构数据转化为标准化表示,支持高效多轮对话,减少信息丢失,提升复杂科学决策能力。为支持训练与评估,我们构建了一个覆盖八个核心药物发现任务的指令数据集。在伯克利函数调用基准测试中,DrugPilot显著优于ReAct和LoT等先进代理,在简单、多工具和多轮场景下的任务完成率分别达到98.0%、93.5%和64.0%,展现出其在需自动化、交互式与数据融合推理的计算科学领域的广泛应用潜力。

原文摘要 · Abstract (English)

Large language models (LLMs) integrated with autonomous agents hold significant potential for advancing scientific discovery through automated reasoning and task execution. However, applying LLM agents to drug discovery is still constrained by challenges such as large-scale multimodal data processing, limited task automation, and poor support for domain-specific tools. To overcome these limitations, we introduce DrugPilot, a LLM-based agent system with a parameterized reasoning architecture designed for end-to-end scientific workflows in drug discovery. DrugPilot enables multi-stage research processes by integrating structured tool use with a novel parameterized memory pool. The memory pool converts heterogeneous data from both public sources and user-defined inputs into standardized representations. This design supports efficient multi-turn dialogue, reduces information loss during data exchange, and enhances complex scientific decision-making. To support training and benchmarking, we construct a drug instruction dataset covering eight core drug discovery tasks. Under the Berkeley function-calling benchmark, DrugPilot significantly outperforms state-of-the-art agents such as ReAct and LoT, achieving task completion rates of 98.0%, 93.5%, and 64.0% for simple, multi-tool, and multi-turn scenarios, respectively. These results highlight DrugPilot's potential as a versatile agent framework for computational science domains requiring automated, interactive, and data-integrated reasoning.

药物发现智能代理大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。