bioMoR利用生物知识提升基因组分析效率,显著降低计算量。
bioMoR: Biology-Guided Mixture-of-Recursions for Effective Genomic Learning

- 将生物结构知识融入递归混合模型,引导基因间注意力与计算分配。
- 在8个基准上平均宏F1提升8.2个百分点,参数减少75%,计算量降58%。
- 适合需要高效、可解释的基因组分析研究者使用。
用于高维组学分析的Transformer模型需处理成千上万的基因或通路,但仅部分需要深度计算。混合递归(MoR)通过自适应选择令牌或专家路由提升效率。我们提出bioMoR,据我们所知是首个将MoR应用于基因级和通路级学习的框架。其贡献包括在MoR主干中识别三种整合结构化生物知识的位置:基于图的信息共享优化令牌嵌入,结构偏差引导自注意力聚焦于生物学相关令牌,图感知路由器利用邻域信息决定每个令牌的递归深度。这些技术基于核心洞察:额外的令牌交互知识有助于模型构建嵌入并选择应深入学习的令牌。在涵盖多种组学数据类型、采用统一五折交叉验证协议的八个基准上,bioMoR相较于最强的无生物学背景MoR基线,平均宏F1提升8.2个百分点,平衡准确率提升7.1个百分点,同时参数减少75%,浮点运算量(FLOPs)最多减少58%,低于非递归Transformer。选中的标志基因或通路提供生物学可解释性,而各令牌特有的递归深度揭示了计算资源的分配情况。
原文摘要 · Abstract (English)
Transformer models for high-dimensional omics analysis process thousands of genes or pathways, although only a subset requires deep computation. Mixture-of-Recursions (MoR) improves efficiency through adaptive token-choice or expert-choice routing. We propose bioMoR, which, to the best of our knowledge, is the first framework to apply MoR to gene-level and pathway-level learning. Our contributions include identifying three locations for integrating structured biological knowledge within an MoR backbone: graph-based information sharing refines token embeddings, a structural bias guides self-attention toward biologically related tokens, and a graph-aware router uses neighborhood information to determine each token's recursion depth. These techniques are centered on our insight that additional knowledge of token interaction can effectively help models construct embeddings and select which tokens should be learned more deeply. Across eight benchmarks spanning diverse omics data types and evaluated under a unified five-fold cross-validation protocol, bioMoR improves average macro-F1 by 8.2 percentage points and balanced accuracy by 7.1 percentage points over the strongest biology-agnostic MoR baseline while using 75 percent fewer parameters and up to 58 percent fewer FLOPs than a non-recursive Transformer. The selected marker genes or pathways provide biological interpretability, while their token-specific recursion depths reveal how computation is allocated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。