不同AI蛋白质折叠模型均通过两阶段机制将氨基酸序列转为三维结构。
Two Stages of Folding: Convergent Mechanisms in AI Protein Folding Trunks
- 先提取电荷等生化信号,再积累距离与接触信息。
- 操控电荷和距离特征可预测性改变蛋白质结构。
- 跨模型的表示可线性对齐,说明底层机制一致。
蛋白质结构预测模型如何折叠蛋白质?我们通过对ESMFold、OpenFold和Boltz-1的折叠主干进行因果干预研究发现,三者均存在共享的两阶段计算结构。第一阶段中,早期模块初始化成对生化信号:电荷等特征通过架构特异路径从序列传递到成对表示。第二阶段中,晚期模块发展成对空间特征:距离与接触信息在成对表示中逐步累积。我们通过因果验证表明,引导电荷与距离特征可引发可预测的结构变化。此外,这些表示具有功能可互换性:成对状态可在不同模型间线性对齐并替换。结果表明,尽管架构、输入和训练方式各异,折叠主干仍收敛于统一的表征组织,实现从序列化学到空间几何的映射。
原文摘要 · Abstract (English)
How do protein structure prediction models fold proteins? We investigate this question through causal interventions on the folding trunks of ESMFold, OpenFold, and Boltz-1. Across all three models, we find a shared two-stage computational structure. In the first stage, early blocks initialize pairwise biochemical signals: features like charge propagate from sequence into pairwise representations through architecture-specific pathways. In the second stage, late blocks develop pairwise spatial features: distance and contact information accumulate in the pairwise representation. We verify these mechanisms causally by showing that steering charge and distance features induces predictable structural changes. Furthermore, these representations are functionally interchangeable: pairwise states can be linearly aligned and substituted across models. Together, these results suggest that folding trunks with different architectures, inputs, and training procedures converge on a shared representational organization for mapping sequence chemistry into spatial geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。