通过新方法解析蛋白质共折叠模型的内部机制,让复杂预测结果可解释。
PairSAE: Mechanistic Interpretability from Pair Representations in Protein Co-Folding

- 将成对表示用N模奇异值分解提炼为残基交互角色
- 在PLINDER数据集上成功识别出与注释一致的可解释特征
- 适合研究生物大分子模型可解释性的研究人员使用
结构生物学领域的基础模型在预测生物分子结构方面表现卓越,并有望用于蛋白质和小分子的设计。然而,理解其输出背后的内部特征仍具挑战性。标准稀疏自编码器(SAEs)虽在Transformer类序列嵌入中有效,但难以直接应用于Pairformer类架构:若直接作用于成对表示,会导致特征数量呈二次增长,并掩盖同时分布在序列与成对表示中的概念。本文提出PairSAE,先通过N模奇异值分解(N-mode SVD)将成对张量压缩为残基级交互角色,再使用稀疏自编码器学习一组共享的残基级特征,这些特征可解码回序列和成对表示。在Boltz-2对PLINDER蛋白-配体复合物的激活值上评估显示,PairSAE生成的特征与UniProt注释高度对齐,并能预测Boltz-2亲和力值。结果表明,PairSAE将结构生物学基础模型的隐空间与可解释的结构概念关联起来,揭示了模型所‘知晓’的内容,同时规避了传统SAEs在Pairformer架构下的固有缺陷。
原文摘要 · Abstract (English)
Foundation models for structural biology have achieved remarkable performance in predicting biomolecular structure and show promise for the design of proteins and small molecules. Yet understanding which internal features drive their outputs remains challenging. Standard sparse autoencoders (SAEs), effective on transformer-style sequence embeddings, do not transfer cleanly to pairformer-like architectures: naively operating on pairwise representations yields a quadratic blow-up of features and obscures concepts distributed jointly across sequence and pair representations. We introduce PairSAE, which summarizes pairwise tensors via an N-mode SVD into token-wise interaction roles, then uses a sparse autoencoder to learn a shared set of token-level features that decode into both sequence and pair representations. Evaluated on Boltz-2 activations for PLINDER protein-ligand complexes, PairSAE yields interpretable features that align with UniProt annotations and predict Boltz-2 affinity values. These results indicate that PairSAE links the latent space of foundation models for structural biology to interpretable structural concepts, clarifying what the model "knows" while avoiding pairformer-induced pitfalls that limit conventional SAEs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。