arXiv:2603.20825cs.LG2026-03

跨粒度表示提升生物序列模型性能与可解释性

Cross-Granularity Representations for Biological Sequences: Insights from ESM and BiGCARP

  • 通过分析不同粒度的嵌入表示,发现深层特征更忠实反映模型知识
  • 不同粒度表示编码互补生物信息,融合后中间层任务表现提升
  • 适合关注生物序列建模可解释性与性能优化的研究者

通用基础模型的发展推动了大规模生物序列模型的兴起。与自然语言的符号粒度(字符、词、句子)不同,生物序列具有层级粒度(核苷酸、氨基酸、蛋白域、基因),其各层级蕴含生物学功能信息。本文以针对生物合成基因簇的Pfam域级模型BiGCARP和氨基酸级蛋白质语言模型ESM为例,通过表征分析工具和一系列探测任务,揭示了直接跨模型嵌入初始化无法提升BiGCARP下游性能的原因,并表明深层嵌入更能捕捉上下文相关的模型知识。进一步证明,不同粒度的表示编码互补的生物学知识,融合后在中等粒度预测任务上取得显著性能提升。研究强调跨粒度整合是提升生物基础模型性能与可解释性的有效策略。

原文摘要 · Abstract (English)

Recent advances in general-purpose foundation models have stimulated the development of large biological sequence models. While natural language shows symbolic granularity (characters, words, sentences), biological sequences exhibit hierarchical granularity whose levels (nucleotides, amino acids, protein domains, genes) further encode biologically functional information. In this paper, we investigate the integration of cross-granularity knowledge from models through a case study of BiGCARP, a Pfam domain-level model for biosynthetic gene clusters, and ESM, an amino acid-level protein language model. Using representation analysis tools and a set of probe tasks, we first explain why a straightforward cross-model embedding initialization fails to improve downstream performance in BiGCARP, and show that deeper-layer embeddings capture a more contextual and faithful representation of the model's learned knowledge. Furthermore, we demonstrate that representations at different granularities encode complementary biological knowledge, and that combining them yields measurable performance gains in intermediate-level prediction tasks. Our findings highlight cross-granularity integration as a promising strategy for improving both the performance and interpretability of biological foundation models.

生物序列跨粒度表示学习模型融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。