解析蛋白质结构模型中信息如何跨模块传递,发现序列化学信息在扩散阶段被显著削弱。
Probing and steering biology across Boltz-1s trunk-diffusion boundary

- 用线性探测和因果干预分析模型各模块的残基激活特征
- 二级结构信息在扩散模块中几乎不变,而氨基酸化学特性明显衰减
- 可解码不等于能控制,部分方向虽可预测却无法有效调控结构
AlphaFold3 类结构预测模型由处理序列与上下文的表征主干(trunk)和生成原子坐标的扩散模块组成。生物信息如何跨越这一架构边界仍不清楚。我们通过线性探测、稀疏自编码器(SAEs)和因果干预,分析 Boltz-1 模型中 Pairformer 主干与扩散模块的残基激活。主干中,几何特征(二级结构、无序性)和序列化学特征(氨基酸身份、信号肽、二硫键注释)均可线性解码。在扩散模块中,两者发生分化:二级结构几乎完全保留,而序列化学特性显著衰减。随后测试这些可解码方向是否能引导模型生成结构,对条件扩散模块的最终主干单表示进行干预。螺旋与卷曲方向在剂量依赖下改变预测结构,优于匹配范数随机对照;但高预测力的β-折叠方向(F1=0.82)未能提升实际折叠含量:线性可解码不等于因果影响。同一探测器在稀疏的 SwissProt 注释上得分显著低于密集的 DSSP 标签,因未标注且模型正确的位置被计为假阳性,故得分是下限。监督式探测器在已有标签处始终优于单一 SAE 特征。本文发布训练好的主干与扩散模块 SAE、Boltz-1 残基激活数据及分析代码。
原文摘要 · Abstract (English)
AlphaFold3-class structure predictors pair a representational trunk, which processes sequence and context, with a diffusion module, which generates atomic coordinates. How biological information changes as it crosses this architectural boundary remains poorly understood. We analyze per-residue activations from the Pairformer trunk and diffusion module of Boltz-1 using linear probes, sparse autoencoders (SAEs), and causal interventions. From the trunk, both geometry (secondary structure, disorder) and sequence chemistry (amino-acid identity, signal peptides, disulfide-bond annotations) are linearly decodable. In the diffusion module, the two diverge. Secondary structure transfers essentially unchanged, whereas sequence chemistry is strongly attenuated. We then test whether decodable directions can steer the model, intervening on the final trunk single representation that conditions the diffusion module. Helix and coil directions change predicted structure dose-dependently against matched-norm random controls, but a beta-strand direction that is highly predictive (F1 =0.82) produces no measurable increase in strand content: linear decodability does not imply causal influence at the site we tested. The same probes also score markedly lower against sparse SwissProt annotations than against dense DSSP labels, because unannotated residues that the model gets right are charged as false positives; such scores are therefore lower bounds. Finally, supervised probes outscore single SAE features wherever a label already exists. We release the trained trunk and diffusion SAEs, Boltz-1 per-residue activations, and the analysis code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。