arXiv:2510.02476cs.LGq-bio.QM2025-10中稿 · NeurIPS被引 2

用模型不确定性选最佳预测器,提升生物分子效力预测精度。

Uncertainty-Guided Model Selection for Tabular Foundation Models in Biomolecule Efficacy Prediction

  • 基于模型预测不确定性的高低筛选最优模型进行集成。
  • 新方法在siRNA抑制效力预测中超越现有顶尖模型。
  • 无需真实标签,适合缺乏标注数据的生物医学预测场景。

上下文学习模型如TabPFN在生物分子效力预测中表现优异,可利用已有的分子特征和实验结果作为上下文示例。然而其性能对上下文高度敏感,因此采用不同数据子集训练多个模型并后处理集成是一种可行策略。本文探讨了一种基于不确定性的模型选择方法:在siRNA敲降效力任务中,仅使用简单序列特征的TabPFN模型即超越了专用的前沿预测器;同时发现模型预测的四分位距(IQR)与真实误差呈负相关。据此提出OligoICP方法,通过选择并平均低均值IQR的模型集合,在siRNA效力预测上优于直接集成或单模型全数据训练。该结果表明模型不确定性可作为无标签情况下的有效优化依据。

原文摘要 · Abstract (English)

In-context learners like TabPFN are promising for biomolecule efficacy prediction, where established molecular feature sets and relevant experimental results can serve as powerful contextual examples. However, their performance is highly sensitive to the provided context, making strategies like post-hoc ensembling of models trained on different data subsets a viable approach. An open question is how to select the best models for the ensemble without access to ground truth labels. In this study, we investigate an uncertainty-guided strategy for model selection. We demonstrate on an siRNA knockdown efficacy task that a TabPFN model using straightforward sequence-based features can surpass specialized state-of-the-art predictors. We also show that the model's predicted inter-quantile range (IQR), a measure of its uncertainty, has a negative correlation with true prediction error. We developed the OligoICP method, which selects and averages an ensemble of models with the lowest mean IQR for siRNA efficacy prediction, achieving superior performance compared to naive ensembling or using a single model trained on all available data. This finding highlights model uncertainty as a powerful, label-free heuristic for optimizing biomolecule efficacy predictions.

模型选择生物医学不确定性表格模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。