arXiv:2608.22642cs.LGcs.AI2026-08

用多模态掩码学习分子表示,提升药物发现效果。

Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules

论文配图:Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules
图 1 · 摘自论文原文
  • 通过掩码多模态数据预测潜在表示,避免化学无效扰动。
  • 在多个基准上表现优异,验证了生物化学上下文的重要性。
  • 适合药物研发与分子表征学习研究者参考。

尽管分子基础模型取得进展,仍存在化学无效增强、模态坍缩及生化环境表征不完整等挑战。为此,我们提出可扩展的分子世界模型框架 Mol-JEPA。该模型不依赖次优的分子扰动,而是通过模态掩码,利用分子结构、细胞表型、结合亲和力、ADMET特征、量子化学模拟及其他药物发现数据中的信息。在多个基准测试中,Mol-JEPA 学习到的表示展现出强大性能,证明了通过潜在空间预测融入生化上下文的价值。

原文摘要 · Abstract (English)

Despite recent advances in molecular foundation models, several limitations remain, such as chemically invalid augmentations, modality collapse, and incomplete representation of biochemical environments. To address these challenges, we present \textbf{Mol-JEPA}, a scalable framework for learning molecular world models. Rather than relying on suboptimal molecular perturbations, our model uses modality masking to exploit information from molecular structures, cellular phenotypes, binding affinities, ADMET profiles, quantum chemistry simulations and other drug discovery data. Across various benchmarks, we show that the representations learned by Mol-JEPA deliver strong performance, demonstrating the value of incorporating biochemical context through latent space prediction.

分子建模多模态生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。