融合蛋白与基因组信息,提升微生物操纵子预测准确性
MicroFuse: Protein-to-Genome Expert Fusion for Microbial Operon Reasoning

- 用四个专家模型分别处理蛋白、基因组、一致和冲突信号,动态融合多模态信息
- 在10万对样本上测试,各项指标均优于单一模型或简单拼接方法
- 特别擅长处理蛋白相似但基因组布局暗示独立调控的复杂情况
预测微生物操纵子共定位需整合两类互补生物信号:蛋白质尺度分子身份与基因组上下文组织。现有生物基础模型虽能独立表征每种视角,但简单拼接会忽略关键生物学特性——当相邻基因构成功能模块时,蛋白身份与基因组布局应一致;而当序列相似性具有误导性但基因组结构显示独立调控时,则应存在冲突。本文提出MicroFuse,一种蛋白到基因组专家融合框架,将ProstT5的结构感知蛋白表示与Bacformer的基因组上下文表示通过四专家混合专家模块(蛋白、基因组上下文、一致、冲突专家)及可学习软路由进行融合。训练结合二元交叉熵、对称跨模态InfoNCE对齐与分歧加权监督对比损失。我们进一步构建了基于OMG宏基因组语料库的OG-Operon100K基准数据集,包含10万对支架级样本,采用生物学合理的正负样本标准。在该数据集上,MicroFuse在所有指标(AUROC、AUPRC、mAP、mAR)上均优于仅用ProstT5、仅用Bacformer及串联MLP基线。消融实验表明跨模态对比对齐是主导因素,硬冲突子集分析显示其在蛋白身份误导但基因组布局明确的生物学模糊情形下表现最优。
原文摘要 · Abstract (English)
Predicting microbial operon co-membership requires integrating two complementary biological signals: protein-scale molecular identity and genome-context organization. While recent biological foundation models provide powerful representations of each view independently, naive concatenation of these modalities ignores a key biological property -- protein identity and genomic context may agree when adjacent genes form a coherent functional module, or conflict when sequence similarity is misleading but genomic layout indicates independent regulation. We present MicroFuse, a protein-to-genome expert fusion framework that integrates structure-aware protein representations from ProstT5 with genome-context representations from Bacformer through a four-expert Mixture-of-Experts module (protein, genome-context, agreement, and conflict experts) with a learned soft router. Training combines binary cross-entropy with symmetric cross-modal InfoNCE alignment and disagreement-weighted supervised contrastive shaping. We further construct OG-Operon100K, a 100,000-pair scaffold-level benchmark from the OMG metagenomic corpus with biologically grounded positive and negative criteria. On OG-Operon100K, MicroFuse achieves the strongest AUROC, AUPRC, mAP, and mAR among ProstT5-only, Bacformer-only, and Concat MLP baselines. Ablations identify cross-modal contrastive alignment as the dominant component, and a hard sequence-conflict subset reveals MicroFuse's largest gains precisely in biologically ambiguous cases where protein identity alone is misleading.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。