arXiv:2601.15333cs.LGcs.AI2026-01被引 2

用隐空间探索增强LLM,提升药物分子设计能力

Empowering LLMs for Structure-Based Drug Design via Exploration-Augmented Latent Inference

  • 将生成过程拆解为编码-隐空间探索-解码,主动扩展模型知识边界
  • 在CrossDocked2020上生成分子结合亲和力得分优于7个基线方法
  • 适合需要高效生成有效药物分子的科研与工业研发人员

大型语言模型具备强大的表示与推理能力,但在结构基础药物设计(SBDD)中的应用受限于对蛋白质结构理解不足及分子生成不可控。为此,我们提出面向大模型的探索增强隐空间推理框架(ELILLM),将生成过程重新定义为编码、隐空间探索与解码的流程。该框架在模型未知区域主动探索,同时利用解码模块处理已知区域,生成化学上有效且合成可行的分子。实现中,贝叶斯优化引导隐向量的系统性探索,位置感知代理模型高效预测结合亲和力分布以指导搜索;知识引导解码减少随机性,有效施加化学有效性约束。在CrossDocked2020基准测试中,ELILLM展现出强可控探索能力与高结合亲和力得分,优于七种基线方法,证明其可有效增强大模型在结构基础药物设计中的能力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) possess strong representation and reasoning capabilities, but their application to structure-based drug design (SBDD) is limited by insufficient understanding of protein structures and unpredictable molecular generation. To address these challenges, we propose Exploration-Augmented Latent Inference for LLMs (ELILLM), a framework that reinterprets the LLM generation process as an encoding, latent space exploration, and decoding workflow. ELILLM explicitly explores portions of the design problem beyond the model's current knowledge while using a decoding module to handle familiar regions, generating chemically valid and synthetically reasonable molecules. In our implementation, Bayesian optimization guides the systematic exploration of latent embeddings, and a position-aware surrogate model efficiently predicts binding affinity distributions to inform the search. Knowledge-guided decoding further reduces randomness and effectively imposes chemical validity constraints. We demonstrate ELILLM on the CrossDocked2020 benchmark, showing strong controlled exploration and high binding affinity scores compared with seven baseline methods. These results demonstrate that ELILLM can effectively enhance LLMs capabilities for SBDD.

药物设计大模型隐空间探索分子生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。