arXiv:2604.04287cs.LGcs.CL2026-04中稿 · ICLR

高熵导致基因组模型学习受限,预测不稳定且难以捕捉序列关系。

Entropy, Disagreement, and the Limits of Foundation Models in Genomics

  • 通过分析模型输出与嵌入,发现基因组序列熵高导致分布趋均
  • 同一训练条件下模型间预测分歧大,嵌入结果不一致
  • 自监督训练难捕捉碱基间关联,限制基础模型能力

基因组领域的基础模型相较于自然语言处理表现参差不齐,其原因尚不明确。本文研究熵作为根本因素对模型从训练数据中学习能力的限制。我们在文本和DNA序列上训练模型集成,并分析其预测、静态嵌入及经验费雪信息流。结果显示,从未见词预测角度看,基因组序列熵高导致输出分布趋近均匀,模型间分歧显著,静态嵌入不稳定,即使模型架构、训练和数据完全一致。进一步发现,DNA训练模型的信息集中于嵌入层,未能有效利用跨标记关系。结果表明,仅依赖序列的自监督训练可能不适用于基因组数据,质疑了当前基因组基础模型训练方法的基本假设。

原文摘要 · Abstract (English)

Foundation models in genomics have shown mixed success compared to their counterparts in natural language processing. Yet, the reasons for their limited effectiveness remain poorly understood. In this work, we investigate the role of entropy as a fundamental factor limiting the capacities of such models to learn from their training data and develop foundational capabilities. We train ensembles of models on text and DNA sequences and analyze their predictions, static embeddings, and empirical Fisher information flow. We show that the high entropy of genomic sequences -- from the point of view of unseen token prediction -- leads to near-uniform output distributions, disagreement across models, and unstable static embeddings, even for models that are matched in architecture, training and data. We then demonstrate that models trained on DNA concentrate Fisher information in embedding layers, seemingly failing to exploit inter-token relationships. Our results suggest that self-supervised training from sequences alone may not be applicable to genomic data, calling into question the assumptions underlying current methodologies for training genomic foundation models.

基因组基础模型熵分析自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。