arXiv:2511.12797cs.LGcs.AI2025-11

基因组模型也能像语言模型一样通过上下文学习,揭示了通用学习机制的存在。

Genomic Next-Token Predictors are In-Context Learners

  • 在基因组序列上训练的模型通过自监督预测实现上下文学习。
  • 上下文示例越多,模型对模式的识别能力呈对数线性提升。
  • 首次证明基因组中可自发产生类语言的上下文学习能力,适合对通用智能感兴趣的读者。

上下文学习(ICL)指模型通过输入中的示例推断并应用抽象规律的能力,此前主要在人类语言大模型中研究。本研究探讨这一能力是否能在其他符号序列领域自发出现。我们以训练于基因组序列的Evo2模型为对象,其规模与中等大小的语言模型相当,任务为预测下一个碱基(A/T/C/G)。构建了语言与基因组形式一致的符号推理任务,实现跨模态对比。结果表明,基因组模型的模式识别能力随上下文示范数量增加而呈现对数线性增长,与语言模型表现一致。这是首次在基因组序列中发现有机涌现的上下文学习现象,支持了‘大尺度预测建模’是引发此类元学习的根本原因。该发现将涌现式元学习从语言扩展至非语言领域,暗示一种跨模态统一的上下文学习机制。

原文摘要 · Abstract (English)

In-context learning (ICL) -- the capacity of a model to infer and apply abstract patterns from examples provided within its input -- has been extensively studied in large language models trained for next-token prediction on human text. In fact, prior work often attributes this emergent behavior to distinctive statistical properties in human language. This raises a fundamental question: can ICL arise organically in other sequence domains purely through large-scale predictive training? To explore this, we turn to genomic sequences, an alternative symbolic domain rich in statistical structure. Specifically, we study the Evo2 genomic model, trained predominantly on next-nucleotide (A/T/C/G) prediction, at a scale comparable to mid-sized LLMs. We develop a controlled experimental framework comprising symbolic reasoning tasks instantiated in both linguistic and genomic forms, enabling direct comparison of ICL across genomic and linguistic models. Our results show that genomic models, like their linguistic counterparts, exhibit log-linear gains in pattern induction as the number of in-context demonstrations increases. To the best of our knowledge, this is the first evidence of organically emergent ICL in genomic sequences, supporting the hypothesis that ICL arises as a consequence of large-scale predictive modeling over rich data. These findings extend emergent meta-learning beyond language, pointing toward a unified, modality-agnostic view of in-context learning.

上下文学习基因组建模大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。