arXiv:2605.16331q-bio.BMcs.AI2026-05

蛋白模型预测起始甲硫氨酸,实为统计检索而非生物识别。

Retrieval and competition: how a protein foundation model starts a protein

论文配图:Retrieval and competition: how a protein foundation model starts a protein
图 1 · 摘自论文原文
  • 通过分层查询整合位置信号,从序列首部检索甲硫氨酸偏好。
  • 在非甲硫氨酸起始的序列上仍错误预测甲硫氨酸,准确率仅72%。
  • 揭示模型依赖统计先验,适合研究可解释性与机制验证的读者。

蛋白质语言模型日益用于指导实验与临床决策,但其高置信度预测究竟是基于生物证据识别,还是统计默认值的检索尚不明确。本文以蛋白质普遍起始于甲硫氨酸这一基本规则为例,追踪ESM2-8M模型生成该预测的计算路径。模型并未在掩码位置直接检测甲硫氨酸,而是通过跨层组装的位置特异性查询,从序列起始标记的参考表示中检索甲硫氨酸偏好信号,并通过上下文依赖电路的竞争输出最终结果。为解析位置信息如何传递至输出端,引入注意力分数在旋转频率带中的范数-方向分解。位置编码通过各频带中查询范数与角度对齐的耦合变化实现。在真实起始位非甲硫氨酸的序列上,模型仍预测甲硫氨酸,正确率仅为72%。这并非意外机制带来的正确结果,而是位置先验检索电路匹配统计平均所致,在生物学偏离统计规律时失效。区分二者需在个体电路、频率带及查询构成层面进行精细分析,表明当生物意义更高时,机制验证将尤为必要且困难。即便对于最简单的生物学规则,模型预测也由分布式计算电路中介,而非直接识别,提示任务复杂度提升将进一步模糊模型置信度与生物证据之间的关系。

原文摘要 · Abstract (English)

Protein language models are increasingly used to guide experimental and clinical decisions, yet it is often unclear whether a confident prediction reflects recognition of biological evidence or retrieval of a statistical default. We examine this distinction for a near-universal biological rule, that proteins begin with methionine, by tracing the computational pathway through which ESM2-8M produces this prediction. The model does not detect methionine at the masked position. Instead, it retrieves a methionine-favouring signal from a reference representation at the beginning-of-sequence token via a position-specific query assembled across layers, with the final output emerging through competition with context-dependent circuits. To understand how positional information reaches the readout, we introduce a norm-direction decomposition of attention scores within rotary frequency bands. Positional encoding operates through coupled changes in query norm and angular alignment distributed across these bands. On sequences whose true N-terminus is not methionine, where the biological question matters, the model predicts methionine anyway. This is not a correct prediction produced by an unexpected mechanism, but the output of a positional-prior retrieval circuit that matches the statistical average and fails where biology diverges from it. Distinguishing the two requires resolution at the level of individual circuits, frequency bands, and query composition, suggesting that mechanistic verification will be necessary, and challenging, for predictions where the biological stakes are higher. Even for the simplest biological rule, the model's prediction is mediated by a distributed computational circuit rather than direct recognition, suggesting that increasing task complexity will further obscure the relationship between model confidence and underlying biological evidence.

蛋白质语言模型可解释性机制验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。