PRISM通过检索蛋白质结构-序列模式,提升逆折叠设计精度。
PRISM: Enhancing Protein Inverse Folding through Fine-Grained Retrieval on Structure-Sequence Multimodal Representations
- 基于结构-序列多模态表示,细粒度检索自然蛋白中的保守片段。
- 在多个基准上实现最优困惑度与氨基酸恢复率,折叠指标全面领先。
- 适合蛋白质工程、生成模型研究者,尤其关注结构约束建模的场景。
设计能折叠成目标三维结构的蛋白质序列(即逆折叠问题),是蛋白质工程的核心挑战。由于序列空间庞大且局部结构约束重要,现有深度学习方法虽有良好恢复率,但缺乏显式机制复用自然蛋白质中保留的细粒度结构-序列模式。为此,我们提出PRISM——一种用于逆折叠的多模态检索增强生成框架。PRISM从已知蛋白质中检索潜在结构-序列基元,并与混合自交叉注意力解码器结合。该框架被建模为潜在变量概率模型,并采用高效近似实现,兼具理论严谨性与实际可扩展性。在多个基准测试(包括CATH-4.2、TS50、TS500、CAMEO 2022及PDB日期划分)上的实验表明,PRISM在细粒度多模态检索方面表现优异,实现了当前最优的困惑度与氨基酸恢复率,同时显著提升了折叠质量指标(RMSD、TM-score、pLDDT)。
原文摘要 · Abstract (English)
Designing protein sequences that fold into a target 3-D structure, termed as the inverse folding problem, is central to protein engineering. However, it remains challenging due to the vast sequence space and the importance of local structural constraints. Existing deep learning approaches achieve strong recovery rates, however, lack explicit mechanisms to reuse fine-grained structure-sequence patterns conserved across natural proteins. To mitigate this, we present PRISM a multimodal retrieval-augmented generation framework for inverse folding. PRISM retrieves fine-grained representations of potential motifs from known proteins and integrates them with a hybrid self-cross attention decoder. PRISM is formulated as a latent-variable probabilistic model and implemented with an efficient approximation, combining theoretical grounding with practical scalability. Experiments across multiple benchmarks, including CATH-4.2, TS50, TS500, CAMEO 2022, and the PDB date split, demonstrate the fine-grained multimodal retrieval efficacy of PRISM in yielding SoTA perplexity and amino acid recovery, while also improving the foldability metrics (RMSD, TM-score, pLDDT).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。