arXiv:2608.24953q-bio.QMcs.LG2026-08

揭示蛋白质表示中高阶互作的隐藏机制,发现编码阶段的互作感知不等于下游高阶功能。

Beyond Tokens: Probing Higher-Order Epistasis in Learned Protein Representations

论文配图:Beyond Tokens: Probing Higher-Order Epistasis in Learned Protein Representations
图 1 · 摘自论文原文
  • 提出ORBIT框架,分离互作存在、表示可及性与功能恢复三个维度。
  • RIT显著提升配对互作可及性(ΔA_tok,2=0.2468),但无更高阶功能优势。
  • 深层模型改善三阶功能恢复与可及性,四阶仍低于零,说明预测指标掩盖表示本质。

蛋白质适应度景观包含非线性互作,突变效应依赖于其他残基。我们引入ORBIT——一种针对互作阶数的基准测试框架,用于分离互作存在性、表示可及性与功能恢复。首先在已知互作阶数的合成景观上验证沃尔什诊断法,随后在实验测得的GB1适应度景观下采用FLIP 2-vs-rest设置进行分析。对比了岭回归、标准MLP、独立词元、非线性独立词元以及残差互作词元化(RIT)方法。在20组配对训练种子下,主比较(两层隐藏层)未发现各架构在FLIP测试R²、三阶或四阶功能恢复、最终层三阶或四阶可及性上有显著差异。然而,RIT在词元阶段显著提升了配对互作可及性(ΔA_tok,2 = 0.2468,d_z = 1.67,Holm校正后p = 1.14 × 10^-5),而下游无明显更高阶优势。预设深度/容量分析显示,更深的MLP提升了FLIP预测、三阶功能恢复与最终层三阶可及性;四阶可及性相较浅层模型有所提升,但绝对保留测试R²仍低于零。因此,ORBIT揭示了传统预测指标所掩盖的表示层面变化,区分了早期互作感知编码与下游非线性能力构建的高阶结构。

原文摘要 · Abstract (English)

Protein fitness landscapes contain nonlinear interactions in which mutation effects depend on other residues. We introduce ORBIT, an Order-Resolved Benchmarking of Interaction Transformations framework that separates interaction presence, representation accessibility, and functional recovery. ORBIT first validates Walsh-based diagnostics on synthetic landscapes with known interaction order, then analyzes the experimentally measured GB1 fitness landscape under the FLIP 2-vs-rest setting. We compare ridge regression, a standard MLP, independent tokens, nonlinear independent tokens, and Residual Interaction Tokenization (RIT). Across 20 paired training seeds, the primary two-hidden-layer comparison found no significant architecture differences in FLIP test R^2, third- or fourth-order functional recovery, or final-layer third- or fourth-order accessibility. However, RIT significantly increased pairwise accessibility at the token stage relative to both independent-token controls (Delta A_tok,2 = 0.2468, d_z = 1.67, Holm-adjusted p = 1.14 x 10^-5), without a detectable downstream higher-order advantage. A pre-specified depth/capacity analysis showed that deeper MLPs improved FLIP prediction, third-order functional recovery, and final-layer third-order accessibility; fourth-order accessibility also improved relative to the shallow MLP but remained below zero in absolute held-out R^2. ORBIT therefore reveals representation-level changes hidden by conventional prediction metrics and distinguishes early interaction-aware encoding from higher-order structure constructed by downstream nonlinear capacity.

蛋白质设计表示学习互作建模深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。