arXiv:2508.10541cs.LGq-bio.QM2025-08

用大模型精准预测过敏原,解决真实场景难题

Driving Accurate Allergen Prediction with Protein Language Models and Generalization-Focused Evaluation

  • 基于1000亿参数蛋白语言模型xTrimoPGLM构建预测框架
  • 在新过敏原识别、同源蛋白区分等任务中全面领先
  • 适合生物医学与药物安全领域研究者使用

过敏原通常是引发免疫反应的蛋白质,构成重大公共卫生挑战。为准确识别过敏原蛋白,我们提出Applm(基于蛋白语言模型的过敏原预测),利用1000亿参数的xTrimoPGLM模型。结果显示,Applm在多种贴近真实复杂场景的任务中持续优于七种先进方法,包括识别训练集中无相似样本的新过敏原、在高序列相似度的同源蛋白中区分过敏原与非过敏原,以及评估仅导致微小序列变化的突变功能影响。分析表明,xTrimoPGLM在万亿级数据上预训练所捕捉的通用蛋白序列特征,是Applm性能的关键。我们还开源了Applm及精心构建的基准数据集,以推动后续研究。

原文摘要 · Abstract (English)

Allergens, typically proteins capable of triggering adverse immune responses, represent a significant public health challenge. To accurately identify allergen proteins, we introduce Applm (Allergen Prediction with Protein Language Models), a computational framework that leverages the 100-billion parameter xTrimoPGLM protein language model. We show that Applm consistently outperforms seven state-of-the-art methods in a diverse set of tasks that closely resemble difficult real-world scenarios. These include identifying novel allergens that lack similar examples in the training set, differentiating between allergens and non-allergens among homologs with high sequence similarity, and assessing functional consequences of mutations that create few changes to the protein sequences. Our analysis confirms that xTrimoPGLM, originally trained on one trillion tokens to capture general protein sequence characteristics, is crucial for Applm's performance by detecting important differences among protein sequences. In addition to providing Applm as open-source software, we also provide our carefully curated benchmark datasets to facilitate future research.

过敏原预测蛋白语言模型生物信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。