用嵌入匹配蒸馏大基因模型,200倍缩小后仍保持顶尖性能。
Distilling Genomic Models for Efficient mRNA Representation Learning via Embedding Matching

- 通过嵌入层匹配实现小模型对大模型的蒸馏
- 模型规模缩小200倍,仍在mRNA-bench上达顶尖水平
- 适合算力有限时高效建模生物序列的场景
大型基因组基础模型近期取得了显著成果和体外翻译能力。然而这些模型参数量迅速超过数十亿,在计算资源受限时运行成本高昂。为解决此问题,我们提出一种蒸馏框架,将最先进的基因组基础模型中的mRNA表示迁移到一个更小的、专用于mRNA序列的小模型中,使模型规模减少200倍。实验表明,基于嵌入的蒸馏优于基于输出概率的蒸馏方法,后者表现不稳定。在mRNA-bench上的基准测试显示,该蒸馏模型在同规模模型中达到当前最佳性能,并能与更大架构在mRNA相关任务上竞争。结果表明,基于嵌入的mRNA序列蒸馏是一种有效的训练策略,为基因组学中类似高效可扩展的序列建模提供了可能,尤其适用于大模型计算不可行的情况。
原文摘要 · Abstract (English)
Large Genomic Foundation Models have recently achieved remarkable results and in-vivo translation capabilities. However these models quickly grow to over a few Billion of parameters and are expensive to run when compute is limited. To overcome this challenge, we present a distillation framework for transferring mRNA representations from a state of the art genomic foundation model into a much smaller model specialized for mRNA sequences, reducing the size by 200-fold. Embedding-level distillation worked better than logit based methods, which we found unstable. Benchmarking on mRNA-bench demonstrates that the distilled model achieves state-of-the-art performance among models of comparable size and competes with larger architectures for mRNA-related tasks. Our results highlight embedding-based distillation of mRNA sequences as an effective training strategy for biological foundation models. This enables similar efficient and scalable sequence modelling in genomics, particularly when large models are computationally challenging or infeasible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。