破解自回归蛋白语言模型生成机制,发现可解释的生物功能电路
Circuit Tracing in Autoregressive Protein Language Models

- 用跨层转换器重构多层生成过程,捕捉跨层计算依赖
- 零样本发现稀疏隐变量电路,准确还原生成分布与功能评分
- 揭示保守序列模式与蛋白适应度景观的生物学意义
蛋白语言模型(pLMs)能生成自然界中未见的新蛋白序列,但其生成机制仍不明确。现有基于稀疏自编码器和转换器的可解释性方法主要针对表示学习模型,无法捕捉自回归生成所需的计算过程。本文提出ProGenMech框架,将跨层转换器(CLTs)扩展至ProGen3——一个用于因果生成和片段填充的稀疏专家混合模型。与逐层方法不同,CLTs利用所有前序层的稀疏隐变量重建每层,实现对层间生成计算的忠实恢复。我们进一步开发零样本电路发现框架,识别负责蛋白生成与适应度预测的稀疏隐电路。在因果生成和零样本适应度估计任务中,ProGenMech在恢复ProGen3的概率分布和功能评分行为上优于局部转换器基线,且在片段填充任务中匹配原模型的生成分布。此外,恢复出的电路揭示了与保守序列模式和蛋白适应度景观相关的生物学有意义的基序与功能区域,为可解释、可控的蛋白生成奠定基础。
原文摘要 · Abstract (English)
Protein language models (pLMs) can generate novel protein sequences with properties beyond those observed in nature, yet the mechanisms underlying protein generation remain poorly understood. Existing mechanistic interpretability methods based on sparse autoencoders and transcoders primarily focus on protein representation learning models and do not capture the computation required for autoregressive generation. Here, we introduce ProGenMech, a mechanistic interpretability framework for generative protein language models that extends cross-layer transcoders (CLTs) to ProGen3, a sparse Mixture-of-Experts model trained for both causal generation and span infilling. Unlike per-layer approaches, CLTs reconstruct each layer using sparse latent variables from all preceding layers, enabling faithful recovery of inter-layer generative computation. We further develop a zero-shot circuit discovery framework to identify sparse latent circuits responsible for protein generation and fitness prediction. In causal generation and zero-shot fitness estimation tasks, ProGenMech outperforms local transcoder baselines in recovering ProGen3's probability distribution and functional scoring behavior, while matching the original model's generative distribution in span infilling tasks. Moreover, the recovered circuits reveal biologically meaningful motifs and functional regions associated with conserved sequence patterns and protein fitness landscapes, establishing a foundation for interpretable and steerable protein generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。