arXiv:2510.03309cs.LGq-bio.BM2025-10

轻量级对比学习实现药物文本精准检索,无需大规模预训练。

Thin Bridges for Drug Text Alignment: Lightweight Contrastive Learning for Target Specific Drug Retrieval

  • 用冻结的单模态编码器+双线性投影构建轻量对齐桥。
  • 在骨架划分数据集上显著提升同靶点药物区分能力。
  • 适合资源有限但需精准药物检索的医药研究场景。

多模态基础模型在药物发现中前景广阔,但现有方法大多依赖重型预训练或大规模多模态语料。本文探索是否可通过轻量级对比桥——在冻结的单模态编码器上添加双线性投影头——实现化学与文本表示的对齐,而无需训练完整多模态模型。基于ChEMBL中的配对机制,通过对比学习目标训练双重线性投影,对齐ECFP4分子指纹与生物医学句子嵌入。为更好处理共享同一治疗靶点的药物,引入困难负例加权和边界损失。在基于骨架划分的评估设置下(要求跨不相交化学核心泛化),结果表明该方法实现了非平凡的跨模态对齐,并显著优于冻结基线,在同靶点药物区分上表现更优。这些结果表明,轻量对比桥可作为大规模多模态预训练的高效替代方案,支持骨架感知的药物文本对齐与靶点特异性检索,适用于精准医疗场景。

原文摘要 · Abstract (English)

Multimodal foundation models hold promise for drug discovery and biomedical applications, but most existing approaches rely on heavy pretraining or large scale multimodal corpora. We investigate whether thin contrastive bridges, lightweight projection heads over frozen unimodal encoders can align chemical and textual representations without training a full multimodal model. Using paired mechanisms from ChEMBL, we align ECFP4 molecular fingerprints with biomedical sentence embeddings through dual linear projections trained with a contrastive objective. To better handle drugs sharing the same therapeutic target, we incorporate hard negative weighting and a margin loss. Evaluation under scaffold based splits, which require generalization across disjoint chemical cores, demonstrates that our approach achieves non-trivial cross modal alignment and substantially improves within target discrimination compared to frozen baselines. These results suggest that thin bridges offer a compute efficient alternative to large scale multimodal pretraining, enabling scaffold aware drug text alignment and target specific retrieval in precision medicine.

药物检索对比学习轻量模型精准医疗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。