arXiv:2602.00663cs.AIcs.LG2026-02被引 1

让大模型根据分子优化过程中的解释信息,更高效地设计新药分子。

SEISMO: Explanation-Aware, Trajectory-Conditioned LLM Agents for Sample-Efficient Molecular Optimisation

  • 利用语言描述、优化轨迹和可解释性反馈,指导大模型生成分子。
  • 在多个药物发现任务中,样本效率显著优于现有方法。
  • 化学家可自然语言干预,保持对优化过程的控制权。

在化学科学中,尤其是制药领域,优化分子以获得期望属性是核心瓶颈,而分子性质评估常依赖成本高、速度慢的实验检测。为解决这一问题,我们提出SEISMO,一种基于大模型的推理时分子优化代理。该方法不再将评估视为黑盒标量输出,而是利用伴随评估结果的自然语言任务描述、完整的优化轨迹,以及事后可解释性方法与子得分分解生成的机器可读反馈,形成显式引导信号。在多种药物发现相关任务中,该方法持续提升样本效率,且随着解释性反馈的丰富,性能增益更大。实际应用中,药物化学家可审查代理的推理过程,并通过自然语言干预生成方向,确保其在优化项目中的主导地位。

原文摘要 · Abstract (English)

Optimizing molecules to achieve desired properties is a central bottleneck across the chemical sciences, particularly in the pharmaceutical industry, where it underlies the discovery of new drugs. Since molecular property evaluation often relies on costly and rate-limited oracles, such as experimental assays, molecular optimization must be highly sample-efficient. To address this, we introduce SEISMO, an LLM agent for inference-time molecular optimisation that turns information routinely available alongside the oracle score, but discarded by existing methods, into an explicit guidance signal. Rather than treating the oracle as a scalar black box, SEISMO conditions each proposal on a natural-language task description, the full optimization trajectory, and machine-readable feedback derived from post-hoc explainability methods and sub-score decompositions. Across a wide range of drug-discovery-relevant tasks, this consistently improves sample efficiency over existing optimisers as well as zero-shot LLM generation, with gains growing as explanatory feedback is enriched. In practice, medicinal chemists can inspect the agent's reasoning and intervene to steer generation in natural language, keeping them central to molecular optimisation projects.

分子优化大模型药物发现可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。