通过领域适配提升分子语言模型搜索效率,让预训练模型更贴合实际应用。
Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries
- 在目标分子库上微调语言模型编码器,解决预训练与实际任务的领域差异。
- 适配后模型在6个虚拟库中显著提升样本效率,部分表现超越指纹基线。
- 适合需要高效虚拟筛选的药物发现与材料设计团队使用。
预训练分子语言模型被广泛用作学习分子结构-性质关系的编码器,但其在预训练域内外的实际适用性尚不明确。本文系统评估了四种分子语言模型在涵盖药物发现、有机材料和催化领域的六个虚拟分子库上的表现。结果表明,原生模型嵌入在不同库中性能差异显著,而分子指纹则提供稳定可靠的基准。由于存在潜在的领域表示不匹配,我们证明显式领域适配能显著提升表示性能。在目标虚拟库结构上微调语言模型编码器可一致提升样本效率,多个适配模型在基准任务中成为最优表现者。研究显示分子表征质量高度依赖目标领域,且显式适配可增强分子基础模型的实用价值。更广泛而言,领域适配分子表示是实现虚拟筛选与自驱动实验室中样本高效适应决策的有前景策略。
原文摘要 · Abstract (English)
Pretrained molecular language models are increasingly used as molecular encoders for learning structure-property relationships. However, their practical suitability for molecular discovery within and beyond their pretraining domain remains unclear. Herein, we systematically benchmark four molecular language models across six virtual molecular libraries spanning drug discovery, organic materials, and catalysis. Native molecular language model embeddings show substantial variation in discovery performance across libraries, whereas molecular fingerprints provide a consistently strong and robust baseline. Consistent with a potential domain-representation mismatch, we show that explicit domain adaptation substantially improves representation performance. Fine-tuning molecular language model encoders on structures from the target virtual library consistently improves sample efficiency, with several adapted encoders emerging as the top-performing representations across the benchmark tasks. These results show that molecular representation quality depends strongly on the target domain and that explicit adaptation can improve the practical utility of molecular foundation models. More broadly, our findings establish domain-adapted molecular representations as a promising strategy for sample-efficient adaptive decision making in virtual screening and self-driving laboratories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。