arXiv:2608.27754q-bio.QMcs.AI2026-08

用新方法高效发现生物AI模型中的可解释特征。

Efficient Auto-Interpretability of AI Models in Biology

  • 分三步验证潜在特征:稳定性、模式识别、可验证描述。
  • 仅需4.4倍少的计算量,就找到超半数可解释特征。
  • 适合关注生物AI可解释性的研究人员使用。

稀疏自编码器(SAEs)等可解释性方法能将生物学领域中的人工智能模型转化为科学发现引擎,揭示其超人类能力的内在机制。但一个潜在特征是否可用,取决于三个条件:是否具有内在一致性、能否被描述、描述是否具备预测能力。这些常被混淆的问题被整合为统一流程。首先,跨种子字典稳定性筛选出值得深入研究的潜在特征;其次,通过异常检测任务判断激活样本是否存在可识别模式;最后,提出候选生物学解释并转化为可检验的假说,在模拟环境中验证。在Boltz-1 Pairformer主干网络上部署该流程,稳定性筛选使每次探索所需潜在特征评估减少约4.4倍,实测成本降低5.2倍,同时恢复超过一半的可解释特征;外部验证显示所揭示的基序显著富集于预期注释。结果还暗示潜在矛盾:跨种子稳定性可能更倾向于选择结构相关特征,而非功能相关特征。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs), and other interpretability methods could turn AI models in Biology and other fields into engines of scientific discovery by explaining the superhuman capabilities of those models. However, a latent is only useful if we know three things: whether it is coherent, whether it can be described, and whether that description has predictive power. These questions are routinely conflated. We assemble them into a single pipeline and report the practical innovations each stage required. First, cross-seed dictionary stability prioritises which latents are worth spending resources to investigate. Second, an intruder-detection task asks whether a latents activating examples share a recognizable pattern. Third, a separate pass proposes a candidate biological description which we convert into falsifiable predictions which can be tested in silico. Deployed on the Boltz-1 Pairformer trunk, stability prioritisation finds interpretable latents using about 4.4 times fewer latent evaluations each, and at 5.2 times lower measured cost, while recovering over half of them, and the external check shows the surfaced motifs are significantly enriched for their claimed annotations. The results also suggest a possible tension: the cross- seed stability might be selecting for some types of features, like structure-related ones, much more than others, such as function-related features.

可解释性生物AI稀疏编码特征挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。