arXiv:2605.01625cs.LG2026-05

PRIME通过多尺度物理图模型,统一学习蛋白质从原子到整体的结构层次信息。

PRIME: Protein Representation via Physics-Informed Multiscale Equivariant Hierarchies

论文配图:PRIME: Protein Representation via Physics-Informed Multiscale Equivariant Hierarchies
图 1 · 摘自论文原文
  • 构建五层物理驱动的多尺度结构图,层级间通过物理规则连接
  • 在折叠分类任务上超越最强基线18.30分,反应类别预测达84.10%新高
  • 自动识别任务相关结构层级,适合蛋白功能预测与结构分析研究者

蛋白质是多尺度物理系统,其功能特性源于原子相互作用到整体折叠拓扑的协同结构组织。现有蛋白质表示学习方法通常仅在单一结构层级运行,或把不同结构信息视为并行模态,未显式建模层级关系。我们提出PRIME(Protein Representation via Physics-Informed Multiscale Equivariant Hierarchies),一个统一框架,将蛋白质建模为涵盖表面、原子、残基、二级结构和蛋白五个物理基础结构图的嵌套体系。相邻层级通过确定性、物理启发的分配算子连接,实现自底向上聚合与自顶向下上下文精炼的双向信息传递。在标准蛋白质表示学习基准测试中表现优异,尤其在折叠分类任务上,相比最强几何GNN基线,在更难的超家族和折叠划分上分别提升13.80和18.30分;在反应类别预测上达到84.10%的最新准确率,超过所有基线(包括ESM)。消融实验表明各结构层级贡献互补且非冗余,自适应交叉注意力分析显示PRIME能自主识别预测时最相关的结构分辨率。代码已开源。

原文摘要 · Abstract (English)

Proteins are inherently multiscale physical systems whose functional properties emerge from coordinated structural organization across multiple spatial resolutions, ranging from atomic interactions to global fold topology. However, existing protein representation learning methods typically operate at a single structural level or treat different sources of structural information as parallel modalities, without explicitly modeling their hierarchical relationships. We introduce PRIME (Protein Representation via Physics-Informed Multiscale Equivariant Hierarchies), a unified framework that models proteins as a nested family of five physically grounded structural graphs spanning surface, atomic, residue, secondary-structure, and protein levels. Adjacent levels are connected through deterministic, physics-informed assignment operators, enabling bidirectional information exchange via bottom-up aggregation and top-down contextual refinement. Experiments on standard protein representation learning benchmarks demonstrate strong and competitive performance across diverse tasks, with particularly notable gains on the Fold Classification benchmark, where PRIME outperforms the strongest geometric GNN baseline by margins of 13.80 and 18.30 points on the harder Superfamily and Fold splits, and achieves a state-of-the-art accuracy of 84.10\% on Reaction Class prediction, surpassing all baseline methods, including ESM. Ablation studies confirm that each structural level contributes complementary and non-redundant information, and adaptive cross-attention analysis reveals that PRIME autonomously identifies the most task-relevant structural resolutions at prediction time. Our source code is publicly available at https://github.com/HySonLab/PRIME

蛋白质表示多尺度建模几何深度学习结构生物信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。