用对比学习连接分子与细胞表型,实现零样本分子识别。
How Molecules Impact Cells: Unlocking Contrastive PhenoMolecular Retrieval
- 通过对比学习构建分子与细胞表型的联合嵌入空间。
- 在零样本条件下,活性分子识别准确率提升8.1倍,达77.33%。
- 适合药物发现中的虚拟表型筛选,尤其适用于不同浓度和骨架的分子。
预测分子对细胞功能的影响是治疗设计的核心挑战。表型组学实验利用基于显微镜的技术捕获细胞形态,为揭示分子对细胞的影响提供了高通量解决方案。本文学习分子结构与显微镜表型组学实验之间的联合潜在空间,采用对比学习对齐配对样本。具体研究了对比式表型分子检索任务,即在无标注情况下,仅凭表型实验识别分子结构。我们评估了多模态学习中表型组学与分子模态面临的关键挑战,如实验批次效应、无效分子扰动及扰动浓度编码。通过(1)单模态预训练的表型组学模型,(2)新颖的样本间相似性感知损失函数,(3)基于分子浓度表示的模型,显著提升了多模态检索性能。据此提出MolPhenix模型。MolPhenix利用预训练表型组学模型,在不同扰动浓度、分子骨架和活性阈值下均表现优异。特别地,零样本分子检索中活性分子的准确率相较此前最优方法提升8.1倍,达到77.33%的top-1%准确率。该成果为机器学习在虚拟表型筛选中的应用打开新路径,显著助力药物发现。
原文摘要 · Abstract (English)
Predicting molecular impact on cellular function is a core challenge in therapeutic design. Phenomic experiments, designed to capture cellular morphology, utilize microscopy based techniques and demonstrate a high throughput solution for uncovering molecular impact on the cell. In this work, we learn a joint latent space between molecular structures and microscopy phenomic experiments, aligning paired samples with contrastive learning. Specifically, we study the problem ofContrastive PhenoMolecular Retrieval, which consists of zero-shot molecular structure identification conditioned on phenomic experiments. We assess challenges in multi-modal learning of phenomics and molecular modalities such as experimental batch effect, inactive molecule perturbations, and encoding perturbation concentration. We demonstrate improved multi-modal learner retrieval through (1) a uni-modal pre-trained phenomics model, (2) a novel inter sample similarity aware loss, and (3) models conditioned on a representation of molecular concentration. Following this recipe, we propose MolPhenix, a molecular phenomics model. MolPhenix leverages a pre-trained phenomics model to demonstrate significant performance gains across perturbation concentrations, molecular scaffolds, and activity thresholds. In particular, we demonstrate an 8.1x improvement in zero shot molecular retrieval of active molecules over the previous state-of-the-art, reaching 77.33% in top-1% accuracy. These results open the door for machine learning to be applied in virtual phenomics screening, which can significantly benefit drug discovery applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。