arXiv:2606.10080cs.LGcs.AI2026-06

用稀疏自编码器识别蛋白质设计中潜在危险特征,提升安全性与可解释性。

VFUSE: Virulent Feature Understanding with Sparse autoEncoders

论文配图:VFUSE: Virulent Feature Understanding with Sparse autoEncoders
图 1 · 摘自论文原文
  • 在扩散-变压器激活上训练稀疏自编码器,提取可解释的危险特征
  • 识别出仅对危险设计响应的单义特征,准确率达AUROC 0.84
  • 适用于蛋白折叠与合成模型的安全审计,适合安全导向研究者

生成模型在蛋白质设计等领域取得显著进展,但其黑箱特性可能生成有害蛋白。本文提出VFUSE(基于稀疏自编码器的毒力特征理解),通过在扩散-变压器激活上训练稀疏自编码器(SAE),对蛋白模型进行危害感知特征审计。应用于RoseTTAFold3和RFDiffusion3这两款开源蛋白折叠与合成模型时发现,将线性探测器在SAE隐空间中训练,能更有效识别危险设计,且不损失模型性能。进一步地,从SAE中识别出仅在危险设计下激活的单义特征,其AUROC达0.84(q < 10^-13)。据我们所知,这是首个在全原子扩散模型上训练的SAE,也是首个针对蛋白设计模型的特征级毒力审计,为安全可解释的蛋白设计铺平道路。

原文摘要 · Abstract (English)

Generative models have shown remarkable progress in a variety of domains such as protein design, but such power enables the opaque generation of hazardous proteins. In this work, we introduce VFUSE (Virulent Feature Understanding with Sparse autoEncoders), a mechanistic interpretability approach that trains SAEs on diffusion-transformer activations to audit protein models for hazard-aware features. We apply VFUSE to RoseTTAFold3 and RFDiffusion3, popular open-weight models for protein folding and synthesis. We find that for certain blocks, linear probes detect hazardous designs significantly better when fit in the SAE latent space over the original model's representations: improving interpretability without sacrificing model performance. Furthermore, we identify monosemantic features from the SAE that fire only on hazardous designs at up to AUROC $0.84$ ($q < 10^{-13}$). To our knowledge this is the first SAE trained on an all-atom diffusion model and the first feature-level virulence audit of a protein design model, paving the way towards safe and interpretable protein design.

蛋白设计可解释性稀疏自编码器安全审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。