arXiv:2510.27671cs.AIcs.LG2025-10被引 2

用结构序列对齐提升药物设计精准度

MolChord: Structure-Sequence Alignment for Protein-Guided Drug Design

  • 融合文本与序列信息,统一建模蛋白和分子结构
  • 在CrossDocked2020上达到当前最佳性能
  • 适合需要精准靶向设计的药物研发人员

基于结构的药物设计(SBDD)将靶点蛋白映射到候选分子配体,是药物发现的核心任务。有效对齐蛋白结构表示与分子表示,并确保生成药物与其药理特性的一致性,仍是关键挑战。为此,我们提出MolChord,整合两项关键技术:(1) 利用NatureLM——一个统一文本、小分子和蛋白质的自回归模型——作为分子生成器,结合基于扩散的结构编码器,对齐蛋白与分子的结构及其文本描述和序列表示(如蛋白质的FASTA和分子的SMILES);(2) 为引导分子向期望性质演化,我们通过整合偏好数据构建属性感知数据集,并采用直接偏好优化(DPO)精炼对齐过程。在CrossDocked2020上的实验结果表明,该方法在关键评估指标上达到当前最优表现,展现出作为实用化SBDD工具的潜力。

原文摘要 · Abstract (English)

Structure-based drug design (SBDD), which maps target proteins to candidate molecular ligands, is a fundamental task in drug discovery. Effectively aligning protein structural representations with molecular representations, and ensuring alignment between generated drugs and their pharmacological properties, remains a critical challenge. To address these challenges, we propose MolChord, which integrates two key techniques: (1) to align protein and molecule structures with their textual descriptions and sequential representations (e.g., FASTA for proteins and SMILES for molecules), we leverage NatureLM, an autoregressive model unifying text, small molecules, and proteins, as the molecule generator, alongside a diffusion-based structure encoder; and (2) to guide molecules toward desired properties, we curate a property-aware dataset by integrating preference data and refine the alignment process using Direct Preference Optimization (DPO). Experimental results on CrossDocked2020 demonstrate that our approach achieves state-of-the-art performance on key evaluation metrics, highlighting its potential as a practical tool for SBDD.

药物设计结构对齐生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。