arXiv:2605.06879cs.LGq-bio.QM2026-05

考虑进化中存活偏差,提升蛋白质功能预测准确率。

Better Protein Function Prediction by Modeling Survivorship Bias

  • 基于突变机制建模存活偏差,区分可产生却未出现的序列
  • 在流感、呼吸道合胞病毒和新冠病毒预测中超越现有方法
  • 适合研究蛋白质功能与进化关系的生物信息学工作者

自然界中的蛋白质序列数据存在存活偏差:我们仅观察到能存活并繁殖的生物体数据,非功能性突变已被自然选择淘汰。因此,预测蛋白质序列是否具有功能通常只能依赖正样本。尽管正样本-未标记(PU)学习框架可应对此问题,但现有方法忽视了塑造序列可观测性的进化过程。例如,一个距离常见蛋白变异仅一步突变的序列,若功能正常应已被观测到;若未被观测,则暗示其可能无功能。相反,那些极难通过突变产生的序列缺失,可能只是从未出现过。因此,这两类缺失序列在训练时应区别对待。本文提出Evo-PU,一种利用核苷酸突变科学理解来建模单物种均匀监控数据中存活偏差的PU学习框架。在三个任务上——预测保留流感与呼吸道合胞病毒(RSV)诱变研究结果,以及预测未来SARS-CoV-2变异——Evo-PU优于标准PU学习、一类分类(OCC)及蛋白质语言模型(PLMs)。在多物种ProteinGym数据集(覆盖不均)上,我们识别出该方法推广的潜力。

原文摘要 · Abstract (English)

Protein sequence data from nature exhibits survivorship bias: we only observe data from those organisms that survive and reproduce, while non-functional protein mutations are eliminated by natural selection. Thus, predicting whether a protein sequence is functional often requires learning from positive examples alone. While positive-unlabeled (PU) learning frameworks offer a generic solution to this problem, existing PU methods ignore the evolutionary processes that shape sequence observability and cause survivorship bias. Consider a sequence that is one mutation away from a commonly-observed protein variant in a well-surveilled organism. If the sequence were functional, it would likely be observed. If it is not observed, this suggests non-functionality. In contrast, sequences that are unlikely to arise through mutation may be missing simply because they never arose. Thus, these two kinds of missing sequences should be treated differently when training models. In this work, we propose Evo-PU, a PU learning framework that uses a scientific understanding of nucleotide mutation to model survivorship bias for well-surveilled single-organism sequence data. On three prediction tasks using single-organism uniform-coverage surveillance data -- predicting results from held-out influenza and respiratory syncytial virus (RSV) mutagenesis studies, and predicting future SARS-CoV-2 variants -- Evo-PU outperforms standard PU learning, one-class classification (OCC), and protein language models (PLMs). On prediction tasks from multi-organism ProteinGym datasets with more heterogeneous surveillance coverage, we identify opportunities to generalize our approach.

蛋白质预测存活偏差PU学习进化建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。