用表格基础模型实现无需微调的分子性质预测,提升效率与准确性。
Tabular foundation models for in-context prediction of molecular properties

- 通过上下文学习直接预测,无需任务特定训练。
- 在30个MoleculeACE任务上最高达成100%胜率,计算成本更低。
- 适合药物发现、化工设计等数据少场景,尤其适合非专家使用。
准确预测分子性质对药物发现、催化和工艺设计至关重要,但实际应用常受限于小样本数据。分子基础模型虽能学习可迁移的分子表示,但通常需特定任务微调、依赖机器学习知识,且难以超越经典基线。表格基础模型(TFMs)提供新范式:通过上下文学习实现预测,无需任务特定训练。我们评估了多种模型在低至中等数据量下的表现,涵盖标准化制药基准与化工数据集。结果表明,结合冻结的分子基础模型表征(如CheMeleon)、RDKit2d与Mordred等紧凑二维描述符,在多个任务上显著优于传统指纹,30个MoleculeACE任务中最高实现100%胜率,同时大幅降低计算开销。分子表征质量是关键决定因素,基础模型嵌入与2D描述符均带来显著性能提升。该方法为实际应用提供了高效精准的替代方案。
原文摘要 · Abstract (English)
Accurate molecular property prediction is central to drug discovery, catalysis, and process design, yet real-world applications are often limited by small datasets. Molecular foundation models provide a promising direction by learning transferable molecular representations; however, they typically involve task-specific fine-tuning, require machine learning expertise, and often fail to outperform classical baselines. Tabular foundation models (TFMs) offer a fundamentally different paradigm: they perform predictions through in-context learning, enabling inference without task-specific training. Here, we evaluate TFMs in the low- to medium-data regime across both standardized pharmaceutical benchmarks and chemical engineering datasets. We evaluate both frozen molecular foundation model representations, as well as classical descriptors and fingerprints. Across the benchmarks, the approach shows excellent predictive performance while reducing computational cost, compared to fine-tuning, with these advantages also transferring to practical engineering data settings. In particular, combining TFMs with CheMeleon embeddings yields up to 100\% win rates on 30 MoleculeACE tasks, while compact RDKit2d and Mordred descriptors provide strong descriptor-based alternatives. Molecular representation emerges as a key determinant in TFM performance, with molecular foundation model embeddings and 2D descriptor sets both providing substantial gains over classic molecular fingerprints on many tasks. These results suggest that in-context learning with TFMs provides a highly accurate and cost-efficient alternative for property prediction in practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。