arXiv:2506.10271q-bio.QMcs.LG2025-06

用非自然序列测试基因语言模型功能理解能力,发现其依赖进化模式而非机制推理。

Evaluating DNA function understanding in genomic language models using evolutionarily implausible sequences

  • 设计新基准Nullsettes,用无进化先例的合成序列评估模型预测失活突变能力。
  • 12个先进模型在非自然序列上表现差,准确率随原始序列似然降低而急剧下降。
  • 适合关注合成生物学、模型可解释性与功能理解的研究者阅读。

基因语言模型(gLMs)有望为合成生物学生成新型功能性DNA序列,但实现该潜力需模型超越进化合理性,真正理解序列如何编码基因表达与调控。我们提出名为Nullsettes的新基准,用于评估模型在缺乏进化先例的合成表达载体中预测体外致死突变(LOF)的能力。对12个顶尖gLMs的测试显示,多数模型无法稳定识别这些强效致死突变。所有模型在原始序列似然值降低时,预测准确率均出现显著下降,表明其高度依赖对进化模式的匹配,而非对基因表达机制的理解。研究揭示了当前gLMs在应对工程化、非天然序列时的根本局限,并强调亟需更注重功能理解的基准与建模方法。

原文摘要 · Abstract (English)

Genomic language models (gLMs) hold promise for generating novel, functional DNA sequences for synthetic biology. However, realizing this potential requires models to go beyond evolutionary plausibility and understand how DNA sequence encodes gene expression and regulation. We introduce a benchmark called Nullsettes, which assesses how well models can predict in silico loss-of-function (LOF) mutations, in synthetic expression cassettes with little evolutionary precedent. Testing 12 state-of-the-art gLMs, we find that most fail to consistently detect these strong LOF mutations. All models show a sharp drop in predictive accuracy as the likelihood assigned to the original (nonmutant) sequence decreases, suggesting that gLMs rely heavily on pattern-matching to their evolutionary prior rather than on any mechanistic understanding of gene expression. Our findings highlight fundamental limitations in how gLMs generalize to engineered, non-natural sequences, and underscore the need for benchmarks and modeling strategies that prioritize functional understanding.

基因语言模型合成生物学功能理解非自然序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。