arXiv:2507.09982cs.CL2025-07被引 1

用基因组信息和文本描述生成有药效的分子,突破数据孤岛

TextOmics-Guided Diffusion for Hit-like Molecular Generation

  • 通过基因组与分子文本的对应关系构建新数据集
  • 生成的分子既化学合法又具生物相关性,零样本表现优异
  • 适合药物发现研究者,尤其关注靶向治疗的团队

具有治疗潜力的类先导分子生成对靶点特异性药物发现至关重要。然而,该领域缺乏异构数据和统一框架来整合多样化的分子表示。为此,我们提出 TextOmics,一个开创性的基准,建立了组学表达与分子文本描述之间的一一对应关系。TextOmics 提供了一个异构数据集,支持通过表示对齐实现分子生成。基于此,我们提出 ToDi 生成框架,联合依赖组学表达和分子文本描述,生成具有生物学相关性、化学合法性及类先导特征的分子。ToDi 采用两个编码器(OmicsEn 与 TextEn)捕捉多层次生物与语义关联,并设计条件扩散(DiffGen)实现可控生成。大量实验验证了 TextOmics 的有效性,并表明 ToDi 超越现有最先进方法,在零样本治疗分子生成方面展现出显著潜力。代码已开源:https://github.com/hala-ToDi。

原文摘要 · Abstract (English)

Hit-like molecular generation with therapeutic potential is essential for target-specific drug discovery. However, the field lacks heterogeneous data and unified frameworks for integrating diverse molecular representations. To bridge this gap, we introduce TextOmics, a pioneering benchmark that establishes one-to-one correspondences between omics expressions and molecular textual descriptions. TextOmics provides a heterogeneous dataset that facilitates molecular generation through representations alignment. Built upon this foundation, we propose ToDi, a generative framework that jointly conditions on omics expressions and molecular textual descriptions to produce biologically relevant, chemically valid, hit-like molecules. ToDi leverages two encoders (OmicsEn and TextEn) to capture multi-level biological and semantic associations, and develops conditional diffusion (DiffGen) for controllable generation. Extensive experiments confirm the effectiveness of TextOmics and demonstrate ToDi outperforms existing state-of-the-art approaches, while also showcasing remarkable potential in zero-shot therapeutic molecular generation. Sources are available at: https://github.com/hala-ToDi.

分子生成扩散模型药物发现多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。