arXiv:2602.17162cs.AIq-bio.GN2026-02被引 8

用隐空间语义对齐提升基因组模型的泛化能力

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

  • 通过联合嵌入预测架构,让模型预测被掩码片段的功能表示
  • 在17个任务上实现线性探测和零样本性能提升
  • 适用于各类基因组基础模型,尤其适合追求功能理解的研究者

基因组基础模型(GFMs)通常依赖掩码语言建模(MLM)或下一个词元预测(NTP)学习自然规律。尽管在捕捉局部语法方面有效,但这些生成范式更关注词元级重建,而非高层次功能上下文。我们提出JEPA-DNA,一种与模型无关的持续训练框架,将联合嵌入预测架构(JEPA)与传统生成目标结合。通过在潜在空间中监督全局序列嵌入,JEPA-DNA迫使模型预测被掩码基因组片段的功能表示,将学习信号从词元恢复转向语义对齐。我们在17个多样化的基因组基准任务上评估了JEPA-DNA,无论底层GFM架构或生成目标如何,均实现了线性探测和零样本性能的稳定提升。该框架为GFMs建立了新基准,超越现有最佳模型,弥合了生成精度与潜在语义接地之间的差距。通过大量消融实验,我们进一步揭示了生成与潜在目标间的协同作用。代码已公开于https://github.com/NVIDIA-Digital-Bio/JEPA-DNA。

原文摘要 · Abstract (English)

Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature". While effective at capturing local syntax, these generative paradigms prioritize token-level reconstruction over high-level functional context. We introduce JEPA-DNA, a model-agnostic continual training framework that integrates a Joint-Embedding Predictive Architecture (JEPA) with traditional generative objectives. By supervising global sequence embeddings in a latent space, JEPA-DNA forces models to predict the functional representations of masked genomic segments, shifting the learning signal from token recovery to semantic alignment. We evaluate JEPA-DNA on 17 diverse genomic benchmark tasks, demonstrating consistent gains in linear probing and zero-shot performance regardless of the underlying GFM architecture or generative objective. Our framework establishes a new state-of-the-art for GFMs, surpassing the best existing models by bridging generative precision with latent semantic grounding. Through extensive ablation studies, we further characterize the synergistic interplay between generative and latent objectives. Our code is publicly available at https://github.com/NVIDIA-Digital-Bio/JEPA-DNA.

基因组模型联合嵌入基础模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。