利用已有实验数据生成新药分子,提升早期药物发现效率。
Assay2Mol: large language model-based drug design using BioAssay context
- 基于大模型检索相似靶点的实验数据,生成候选分子
- 在生成可合成分子方面优于现有机器学习方法
- 适合需要快速筛选化合物的药物研发团队
科学数据库中积累了大量定量数据和描述性文本。在生物化学领域,分子筛选实验评估候选分子对疾病靶点的功能响应。描述靶点作用机制、实验筛选流程及其他实验属性的非结构化文本,为药物发现提供了丰富信息,但因格式不统一而长期未被利用。我们提出 Assay2Mol,一种基于大语言模型的工作流,可挖掘现有生化筛选实验数据以支持早期药物发现。Assay2Mol 通过检索与新靶点相似的实验记录,并结合上下文学习生成候选分子。该方法在生成候选配体方面优于近期基于蛋白结构的机器学习方法,同时显著提升了生成分子的可合成性。
原文摘要 · Abstract (English)
Scientific databases aggregate vast amounts of quantitative data alongside descriptive text. In biochemistry, molecule screening assays evaluate candidate molecules' functional responses against disease targets. Unstructured text that describes the biological mechanisms through which these targets operate, experimental screening protocols, and other attributes of assays offer rich information for drug discovery campaigns but has been untapped because of that unstructured format. We present Assay2Mol, a large language model-based workflow that can capitalize on the vast existing biochemical screening assays for early-stage drug discovery. Assay2Mol retrieves existing assay records involving targets similar to the new target and generates candidate molecules using in-context learning with the retrieved assay screening data. Assay2Mol outperforms recent machine learning approaches that generate candidate ligand molecules for target protein structures, while also promoting more synthesizable molecule generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。