arXiv:2602.04696physics.chem-phcs.LG2026-02

用廉价的分子基团描述训练模型,让分子表征自动适配任务。

Beyond Learning on Molecules by Weakly Supervising on Molecules

  • 基于程序生成的分子基团和自然语言描述进行弱监督学习
  • 在多个分子性质预测任务上达到顶尖性能,且表征可解释
  • 适合需要快速适配新任务的药物研发与分子设计场景

分子表征本质上依赖具体任务,但现有预训练分子编码器通常不具备这种特性。任务条件化表征可通过任务描述重新组织表示,但现有方法依赖昂贵的标注数据。本文证明,仅需对程序生成的数百个分子基团进行弱监督即可。我们的自适应化学嵌入模型(ACE-Mol)利用低成本计算、易扩展的自然语言描述进行训练。传统编码器需缓慢搜索嵌入空间以寻找任务相关结构,而ACE-Mol能立即对齐表征与任务需求。ACE-Mol在多个分子性质预测基准测试中表现卓越,且生成的表征具有化学意义与可解释性。

原文摘要 · Abstract (English)

Molecular representations are inherently task-dependent, yet most pre-trained molecular encoders are not. Task conditioning promises representations that reorganize based on task descriptions, but existing approaches rely on expensive labeled data. We show that weak supervision on programmatically derived molecular motifs is sufficient. Our Adaptive Chemical Embedding Model (ACE-Mol) learns from hundreds of motifs paired with natural language descriptors that are cheap to compute, trivial to scale. Conventional encoders slowly search the embedding space for task-relevant structure, whereas ACE-Mol immediately aligns its representations with the task. ACE-Mol achieves state-of-the-art performance across molecular property prediction benchmarks with interpretable, chemically meaningful representations.

分子表征弱监督可解释性药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。