融合图、图像和文本多视角表示,提升药物靶点与性质预测性能。
Multi-view biomedical foundation models for molecule-target and property prediction
- 构建多视图分子嵌入模型,整合图、图像、文本三类分子表示
- 在120多个任务上表现稳定,媲美最优单视图模型
- 可用于阿尔茨海默病相关靶点筛选,发现强结合化合物
高质量的分子表征是生物医学领域基础模型发展的关键。以往工作通常聚焦单一分子表示方式,可能在特定任务上存在优劣。我们提出多视图分子嵌入与后融合方法(MMELON),在基础模型框架下集成图、图像和文本三种分子视图,可轻松扩展至其他表示形式。单视图基础模型分别在最多2亿个分子的数据集上进行预训练。多视图模型在超过120项任务中表现稳健,性能达到最优单视图模型水平,涵盖分子溶解度、ADME性质及对G蛋白偶联受体(GPCRs)的活性预测。我们识别出33个与阿尔茨海默病相关的GPCRs,利用多视图模型从化合物筛选中选出强结合剂,并通过基于结构的建模和关键结合基序识别验证预测结果。
原文摘要 · Abstract (English)
Quality molecular representations are key to foundation model development in bio-medical research. Previous efforts have typically focused on a single representation or molecular view, which may have strengths or weaknesses on a given task. We develop Multi-view Molecular Embedding with Late Fusion (MMELON), an approach that integrates graph, image and text views in a foundation model setting and may be readily extended to additional representations. Single-view foundation models are each pre-trained on a dataset of up to 200M molecules. The multi-view model performs robustly, matching the performance of the highest-ranked single-view. It is validated on over 120 tasks, including molecular solubility, ADME properties, and activity against G Protein-Coupled receptors (GPCRs). We identify 33 GPCRs that are related to Alzheimer's disease and employ the multi-view model to select strong binders from a compound screen. Predictions are validated through structure-based modeling and identification of key binding motifs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。