用可解释机器学习快速分类病毒样颗粒的蛋白构成
Classifying the Stoichiometry of Virus-like Particles with Interpretable Machine Learning
- 基于线性模型构建数据驱动分类流程
- 准确识别影响组装的关键蛋白序列特征
- 适合疫苗研发与结构生物学研究者使用
病毒样颗粒(VLPs)因其激发免疫反应的特性,在疫苗开发中具有重要价值。理解其化学计量比(即形成VLP所需的蛋白亚基数量)对优化疫苗设计至关重要。然而,现有实验方法测定化学计量比耗时且需高纯度蛋白。为此,我们构建了一个新数据集,并提出一种可解释的数据驱动分析流程,采用线性机器学习模型进行分类。我们还探讨了特征编码对模型性能与可解释性的影响,以及识别关键蛋白序列特征的方法。评估结果表明,该流程不仅能有效分类化学计量类型,还能揭示可能影响VLP组装的蛋白特征。本研究使用的数据与代码已公开于 https://github.com/Shef-AIRE/StoicIML。
原文摘要 · Abstract (English)
Virus-like particles (VLPs) are valuable for vaccine development due to their immune-triggering properties. Understanding their stoichiometry, the number of protein subunits to form a VLP, is critical for vaccine optimisation. However, current experimental methods to determine stoichiometry are time-consuming and require highly purified proteins. To efficiently classify stoichiometry classes in proteins, we curate a new dataset and propose an interpretable, data-driven pipeline leveraging linear machine learning models. We also explore the impact of feature encoding on model performance and interpretability, as well as methods to identify key protein sequence features influencing classification. The evaluation of our pipeline demonstrates that it can classify stoichiometry while revealing protein features that possibly influence VLP assembly. The data and code used in this work are publicly available at https://github.com/Shef-AIRE/StoicIML.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。