arXiv:2603.03655cs.AI2026-03被引 5

让大模型在药物研发中可靠执行长周期任务,避免错误累积。

Mozi: Governed Autonomy for Drug Discovery LLM Agents

  • 双层架构:控制层管工具使用,工作层用可组合技能图规划研发流程。
  • 在PharmaBench上优于基线,能生成符合毒性筛选的候选分子。
  • 适合需要高可靠性、强可追溯性的药物研发团队使用。

工具增强型大语言模型代理有望融合科学推理与计算,但在药物研发等高风险领域受限于两大瓶颈:工具使用缺乏管控和长期任务可靠性差。在依赖关系复杂的制药流程中,自主代理常因早期幻觉导致轨迹不可复现,错误逐级放大。为此,我们提出Mozi,一种双层架构:第A层(控制平面)建立受控的监督-执行层级,实现基于角色的工具隔离、限制动作空间,并支持反思式重规划;第B层(工作流平面)将靶点发现到先导优化等阶段转化为有状态、可组合的技能图,结合严格数据契约和关键的人机协同检查点,保障高不确定性决策处的科学有效性。遵循‘自由推理用于安全任务,结构化执行用于长周期流程’的设计原则,Mozi具备内置鲁棒性与逐级可审计性,完全消除误差累积。我们在PharmaBench基准上评估,显示其在任务编排精度上优于现有方法。通过端到端治疗案例研究,证明其可在巨大化学空间中导航,强制执行严格毒性过滤,并生成极具竞争力的体外候选分子,使大模型从脆弱对话者转变为可靠、受控的科研协作者。

原文摘要 · Abstract (English)

Tool-augmented large language model (LLM) agents promise to unify scientific reasoning with computation, yet their deployment in high-stakes domains like drug discovery is bottlenecked by two critical barriers: unconstrained tool-use governance and poor long-horizon reliability. In dependency-heavy pharmaceutical pipelines, autonomous agents often drift into irreproducible trajectories, where early-stage hallucinations multiplicatively compound into downstream failures. To overcome this, we present Mozi, a dual-layer architecture that bridges the flexibility of generative AI with the deterministic rigor of computational biology. Layer A (Control Plane) establishes a governed supervisor--worker hierarchy that enforces role-based tool isolation, limits execution to constrained action spaces, and drives reflection-based replanning. Layer B (Workflow Plane) operationalizes canonical drug discovery stages -- from Target Identification to Lead Optimization -- as stateful, composable skill graphs. This layer integrates strict data contracts and strategic human-in-the-loop (HITL) checkpoints to safeguard scientific validity at high-uncertainty decision boundaries. Operating on the design principle of ``free-form reasoning for safe tasks, structured execution for long-horizon pipelines,'' Mozi provides built-in robustness mechanisms and trace-level audibility to completely mitigate error accumulation. We evaluate Mozi on PharmaBench, a curated benchmark for biomedical agents, demonstrating superior orchestration accuracy over existing baselines. Furthermore, through end-to-end therapeutic case studies, we demonstrate Mozi's ability to navigate massive chemical spaces, enforce stringent toxicity filters, and generate highly competitive in silico candidates, effectively transforming the LLM from a fragile conversationalist into a reliable, governed co-scientist.

药物发现大模型代理自动化研发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。