arXiv:2606.31126cs.LGq-bio.QM2026-06被引 1

用表格模型预测生物分子性质,搭配强大表示效果好。

Can Tabular In-Context Learners Generalize to Biomolecular Property Prediction?

  • 用表格型上下文学习模型+蛋白质/分子表示做少样本预测。
  • 在蛋白适应度回归任务上超越或媲美当前最佳结果。
  • 适合缺乏标注数据的生物分子设计场景,无需定制模型。

从少量标注数据预测生物分子性质是蛋白质工程与小分子设计的核心挑战。随着强预训练编码器提供丰富的固定长度表示,难点已转向构建高效的数据预测器以应对少样本情形。表格式基础模型如TabPFN和TabICL是潜在候选:它们在源自随机因果图的合成表格上预训练,具有非生物相关的因果归纳偏置。尽管这种偏置看似不适用于蛋白质序列或分子图生成过程,但我们发现其仍能有效迁移。将每种方法视为预测器-表示对,在两个领域进行评估。在蛋白适应度回归任务中,这些模型结合ESM Cambrian表示在ProteinGym上达到或超过当前最优表现,并在多样酯酶催化活性数据集上优于特定任务的监督回归器。对于使用ECFP/RDKit描述符的小分子分类任务,在TDC ADMET、MoleculeNet、FS-Mol和DrugOOD四个数据集上虽无单一组合主导,但整体表现与现有任务专用最优模型相当。关键在于,无论蛋白还是小分子少样本任务,这些预测器-表示对均展现出优异性能。结论是:表格式基础模型可成为强生物分子预测器,但必须与表达力强的表示配合使用。

原文摘要 · Abstract (English)

Predicting biomolecular properties from limited labeled data is a central bottleneck in protein engineering and small-molecule design. As strong pretrained encoders now supply rich fixed-length representations, the difficulty has shifted from representation learning to building a data-efficient predictor for the few-shot regime. Tabular foundation models such as TabPFN and TabICL are unlikely candidates for this role: they are in-context learners pretrained on synthetic tables drawn from random causal graphs, a generative prior with no obvious correspondence to the processes that produce protein sequences or molecular graphs. That this tabular, causal inductive bias should transfer to biomolecular data at all is counter-intuitive, yet we find it does. Treating each method as a predictor-representation pair, we evaluate across two domains. We find that on protein fitness regression tasks these in-context learning models coupled with ESM Cambrian representations achieve or exceed state-of-the-art results on ProteinGym, and outperform task-specific supervised regressors on a diverse esterase catalytic activity dataset. For small-molecule classification with ECFP/RDKit descriptors, no single predictor-representation pairing dominates across TDC ADMET, MoleculeNet, FS-Mol, and DrugOOD, but they are competitive with the existing task-specific state-of-the-art. Crucially, on both protein and small-molecule few-shot tasks, these predictor-representation pairs offer strong performance. We conclude that tabular foundation models can be strong biomolecular predictors, but only when coupled with expressive representations.

少样本学习生物分子表格模型表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。