用扩散模型生成结构引导的新型pMHC-I肽库,打破数据偏见。
Generation of structure-guided pMHC-I libraries using Diffusion Models
- 基于晶体结构距离条件生成新肽段,避免实验数据偏见。
- 新肽库覆盖20种高优先级HLA等位基因,保留锚定残基偏好。
- 现有预测模型在此类新设计上表现差,揭示其局限性。
个性化疫苗和T细胞免疫疗法依赖于识别能引发强免疫反应的肽-MHC I类(pMHC-I)相互作用。然而,当前基准和模型继承了质谱与结合实验数据集中的偏差,限制了新肽配体的发现。为此,我们引入了一个基于扩散模型、以晶体结构相互作用距离为条件的结构引导型pMHC-I肽基准数据集。该数据集涵盖20种高优先级HLA等位基因,独立于已知肽段,但仍再现了经典的锚定残基偏好,表明其具备结构泛化能力且无实验数据偏差。利用此资源,我们发现当前最先进的序列基预测模型在识别这些结构稳定设计的结合潜力时表现不佳,揭示其在特定等位基因上的隐性局限。我们的几何感知设计流程生成的肽段具有高预测结构完整性及更高残基多样性,是无偏模型训练与评估的关键资源。代码与数据见:https://github.com/sermare/struct-mhc-dev。
原文摘要 · Abstract (English)
Personalized vaccines and T-cell immunotherapies depend critically on identifying peptide-MHC class I (pMHC-I) interactions capable of eliciting potent immune responses. However, current benchmarks and models inherit biases present in mass-spectrometry and binding-assay datasets, limiting discovery of novel peptide ligands. To address this issue, we introduce a structure-guided benchmark of pMHC-I peptides designed using diffusion models conditioned on crystal structure interaction distances. Spanning twenty high-priority HLA alleles, this benchmark is independent of previously characterized peptides yet reproduces canonical anchor residue preferences, indicating structural generalization without experimental dataset bias. Using this resource, we demonstrate that state-of-the-art sequence-based predictors perform poorly at recognizing the binding potential of these structurally stable designs, indicating allele-specific limitations invisible in conventional evaluations. Our geometry-aware design pipeline yields peptides with high predicted structural integrity and higher residue diversity than existing datasets, representing a key resource for unbiased model training and evaluation. Our code, and data are available at: https://github.com/sermare/struct-mhc-dev.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。