蛋白语言模型通过注意力机制识别重复序列,融合生物知识与语言模式。
Induction Meets Biology: Mechanisms of Repeat Detection in Protein Language Models
- 用注意力机制结合氨基酸相似性,分两阶段检测重复序列
- 诱导头关注对齐片段,提升预测准确率
- 揭示模型如何融合生物学知识与语言建模能力
蛋白质序列中普遍存在重复片段,包括完全相同和带有突变的近似重复。这些重复对蛋白质结构与功能至关重要,推动了数十年的算法研究。近期研究表明,蛋白语言模型(PLMs)可通过掩码词预测行为识别重复序列。为揭示其内部机制,我们研究了模型对精确重复与近似重复的检测方式。结果发现,近似重复的检测机制在功能上包含精确重复的机制。我们进一步刻画该机制,发现两个主要阶段:首先,模型通过通用位置注意力头与生物特化组件(如编码氨基酸相似性的神经元)构建特征表示;随后,诱导头关注重复片段中的对齐标记,促进正确预测。结果表明,PLMs通过结合基于语言的模式匹配与生物特化知识来解决这一生物学任务,为研究更复杂的进化过程提供了基础。
原文摘要 · Abstract (English)
Protein sequences are abundant in repeating segments, both as exact copies and as approximate segments with mutations. These repeats are important for protein structure and function, motivating decades of algorithmic work on repeat identification. Recent work has shown that protein language models (PLMs) identify repeats, by examining their behavior in masked-token prediction. To elucidate their internal mechanisms, we investigate how PLMs detect both exact and approximate repeats. We find that the mechanism for approximate repeats functionally subsumes that of exact repeats. We then characterize this mechanism, revealing two main stages: PLMs first build feature representations using both general positional attention heads and biologically specialized components, such as neurons that encode amino-acid similarity. Then, induction heads attend to aligned tokens across repeated segments, promoting the correct answer. Our results reveal how PLMs solve this biological task by combining language-based pattern matching with specialized biological knowledge, thereby establishing a basis for studying more complex evolutionary processes in PLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。