arXiv:2504.17068cs.LGq-bio.BM2025-04被引 2

语言模型在生物序列中误将重复片段当高适配度,导致预测失真。

In-Context Learning can distort the relationship between sequence likelihoods and biological fitness

  • 利用重复片段作为参照,模型错误提升其似然得分
  • 变压器架构模型对重复序列异常敏感,得分虚高
  • 该现象影响蛋白质与RNA结构预测,需警惕误判

语言模型已成为生物序列可行性的强大预测工具。训练过程中,模型学习氨基酸或核苷酸序列的语法规则。训练后,输入序列可输出似然分数,分数越高表示越符合语法,与实验测得的适应度相关。然而我们发现,在上下文学习中,似然分数与适应度之间的关系会被扭曲。尤其表现为含有重复基序的序列获得异常高的似然分数。我们在多种架构的蛋白语言模型上进行实验,均采用掩码语言建模目标训练,发现变压器模型对此效应尤为敏感。该现象由一种查找操作驱动:模型通过其他重复拷贝定位被掩码位置的身份,这种检索行为会覆盖模型已学先验。此问题在不完全重复序列中仍存在,并延伸至其他生物相关特征,如折叠成发夹结构的反向互补基序。

原文摘要 · Abstract (English)

Language models have emerged as powerful predictors of the viability of biological sequences. During training these models learn the rules of the grammar obeyed by sequences of amino acids or nucleotides. Once trained, these models can take a sequence as input and produce a likelihood score as an output; a higher likelihood implies adherence to the learned grammar and correlates with experimental fitness measurements. Here we show that in-context learning can distort the relationship between fitness and likelihood scores of sequences. This phenomenon most prominently manifests as anomalously high likelihood scores for sequences that contain repeated motifs. We use protein language models with different architectures trained on the masked language modeling objective for our experiments, and find transformer-based models to be particularly vulnerable to this effect. This behavior is mediated by a look-up operation where the model seeks the identity of the masked position by using the other copy of the repeated motif as a reference. This retrieval behavior can override the model's learned priors. This phenomenon persists for imperfectly repeated sequences, and extends to other kinds of biologically relevant features such as reversed complement motifs in RNA sequences that fold into hairpin structures.

语言模型生物序列似然偏差重复基序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。