arXiv:2603.28575cs.LGcs.AI2026-03

用对比学习统一有机与无机抗癌化合物表示,提升跨域药物发现能力

ChemCLIP: Bridging Organic and Inorganic Anticancer Compounds Through Contrastive Learning

  • 基于活性而非结构相似性,构建双编码器对比学习框架
  • 在60种癌细胞线上实现0.899的平均对齐率和超0.85的分类准确率
  • 适用于跨化学领域药物研发,尤其适合缺乏数据的金属药物

抗癌药物研发长期将有机小分子与金属配位复合物视为独立化学领域,尽管二者具有共同的生物学目标。这一差异在数据层面尤为显著:有机化合物有大量筛选数据库,而金属复合物仅有数千个被表征。本文提出ChemCLIP,一种双编码器对比学习框架,通过共享抗癌活性而非结构相似性,学习统一的分子表示。我们整合了44,854个独特有机化合物和5,164个独特金属复合物,标准化于60种癌细胞系。通过带活性感知硬负样本挖掘的并行编码器训练,将结构迥异的化合物映射至256维共享嵌入空间,生物活性相似者聚类在一起。系统评估了四种分子编码策略:Morgan指纹、ChemBERTa、MolFormer和Chemprop,结果显示Morgan指纹表现最优,平均对齐比率达0.899,下游分类任务的AUC分别为0.859(无机)和0.817(有机)。该研究确立对比学习在统一异质化学领域中的有效性,并为多模态化学应用中的编码器选择提供实证指导,其意义可扩展至任何需跨域化学知识迁移的场景。

原文摘要 · Abstract (English)

The discovery of anticancer therapeutics has traditionally treated organic small molecules and metal-based coordination complexes as separate chemical domains, limiting knowledge transfer despite their shared biological objectives. This disparity is particularly pronounced in available data, with extensive screening databases for organic compounds compared to only a few thousand characterized metal complexes. Here, we introduce ChemCLIP, a dual-encoder contrastive learning framework that bridges this organic-inorganic divide by learning unified representations based on shared anticancer activities rather than structural similarity. We compiled complementary datasets comprising 44,854 unique organic compounds and 5,164 unique metal complexes, standardized across 60 cancer cell lines. By training parallel encoders with activity-aware hard negative mining, we mapped structurally distinct compounds into a shared 256-dimensional embedding space where biologically similar compounds cluster together regardless of chemical class. We systematically evaluated four molecular encoding strategies: Morgan fingerprints, ChemBERTa, MolFormer, and Chemprop, through quantitative alignment metrics, embedding visualizations, and downstream classification tasks. Morgan fingerprints achieved superior performance with an average alignment ratio of 0.899 and downstream classification AUCs of 0.859 (inorganic) and 0.817 (organic). This work establishes contrastive learning as an effective strategy for unifying disparate chemical domains and provides empirical guidance for encoder selection in multi-modal chemistry applications, with implications extending beyond anticancer drug discovery to any scenario requiring cross-domain chemical knowledge transfer.

药物发现对比学习化学信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。