让语言模型更准地发现药物,靠的是精准诊断和精简记忆。
Constraint-Aware Corrective Memory for Language-Based Drug Discovery Agents
- 用多模态证据分析任务要求,定位失败原因并生成修正提示。
- 相比顶尖基线,成功找到目标药物的概率提升36.4%。
- 适合需要高可靠性、复杂约束的自动化药物研发场景。
大语言模型使自主药物发现成为可能,但成功与否取决于最终候选分子集是否满足整体协议要求,如集合大小、多样性、结合质量与可开发性。这带来根本性控制难题:代理逐步规划,而任务有效性在整体集合层面判定。现有系统依赖冗长历史和模糊自省,导致故障定位不精确,且代理状态噪声累积。本文提出CACM(Constraint-Aware Corrective Memory)框架,围绕集级诊断与简洁记忆写回机制构建。CACM引入协议审计与实体诊断器,联合分析跨越任务要求、靶点环境与候选集证据的多模态信息,定位协议违规,生成可操作的修复提示,并引导下一步行动偏向最相关修正。为保持规划上下文紧凑,CACM将记忆分为静态、动态与修正通道,并压缩后写回,既保留持久任务信息,又仅暴露最关键的决策失败。实验表明,CACM相比当前最优基线,目标成功率提升36.4%。结果表明,可靠的语言驱动药物发现不仅依赖更强分子工具,更需精准诊断与经济高效的代理状态。
原文摘要 · Abstract (English)
Large language models are making autonomous drug discovery agents increasingly feasible, but reliable success in this setting is not determined by any single action or molecule. It is determined by whether the final returned set jointly satisfies protocol-level requirements such as set size, diversity, binding quality, and developability. This creates a fundamental control problem: the agent plans step by step, while task validity is decided at the level of the whole candidate set. Existing language-based drug discovery systems therefore tend to rely on long raw history and under-specified self-reflection, making failure localization imprecise and planner-facing agent states increasingly noisy. We present CACM (Constraint-Aware Corrective Memory), a language-based drug discovery framework built around precise set-level diagnosis and a concise memory write-back mechanism. CACM introduces protocol auditing and a grounded diagnostician, which jointly analyze multimodal evidence spanning task requirements, pocket context, and candidate-set evidence to localize protocol violations, generate actionable remediation hints, and bias the next action toward the most relevant correction. To keep planning context compact, CACM organizes memory into static, dynamic, and corrective channels and compresses them before write-back, thereby preserving persistent task information while exposing only the most decision-relevant failures. Our experimental results show that CACM improves the target-level success rate by 36.4% over the state-of-the-art baseline. The results show that reliable language-based drug discovery benefits not only from more powerful molecular tools, but also from more precise diagnosis and more economical agent states.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。