arXiv:2504.09060cs.LGcs.AI2025-04NeurIPS

首个融合三维基因组与表观遗传数据的预训练模型,提升多模态解析能力。

Multimodal 3D Genome Pre-training

  • 设计跨模态交互与映射模块,实现结构与表观数据统一表征
  • 构建超百万对样本的数据集,支持高质量预训练
  • 在多个下游任务中超越现有方法,适合基因组功能研究者

深度学习推动了计算生物学中3D基因组分析的进展,但对3D基因组知识的整体理解仍不充分。本文提出MIX-HIC,首个整合3D基因组结构与表观遗传数据的多模态基础模型,获得统一且全面的语义表示。为实现精确的异构语义融合,设计了跨模态交互与映射模块,有效聚合3D基因组知识。此外,引入首个包含超过100万对样本的大型数据集,涵盖Hi-C接触图谱与表观遗传轨道,支持高质量预训练,促进3D基因组功能机制的探索。大量实验表明,MIX-HIC在多种下游任务中显著优于现有最先进方法。本工作为推进3D基因组研究提供了宝贵资源。

原文摘要 · Abstract (English)

Deep learning techniques have driven significant progress in various analytical tasks within 3D genomics in computational biology. However, a holistic understanding of 3D genomics knowledge remains underexplored. Here, we propose MIX-HIC, the first multimodal foundation model of 3D genome that integrates both 3D genome structure and epigenomic tracks, which obtains unified and comprehensive semantics. For accurate heterogeneous semantic fusion, we design the cross-modal interaction and mapping blocks for robust unified representation, yielding the accurate aggregation of 3D genome knowledge. Besides, we introduce the first large-scale dataset comprising over 1 million pairwise samples of Hi-C contact maps and epigenomic tracks for high-quality pre-training, enabling the exploration of functional implications in 3D genomics. Extensive experiments show that MIX-HIC can significantly surpass existing state-of-the-art methods in diverse downstream tasks. This work provides a valuable resource for advancing 3D genomics research.

3D基因组多模态预训练表观遗传

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。