arXiv:2607.19086cs.CV2026-07KDD

提出轻量级多模态融合框架,提升医疗数据跨模态建模效率与泛化能力。

Advancing Multimodal Fusion on Heterogeneous Medical Data with Hybrid Geometry Attention

论文配图:Advancing Multimodal Fusion on Heterogeneous Medical Data with Hybrid Geometry Attention
图 1 · 摘自论文原文
  • 设计分层混合几何注意力模块,逐步融合影像、临床、组学等异构医疗数据
  • 在16个公开数据集上性能提升最高3.97%,计算开销降低87.8%
  • 支持灵活模态组合,适合资源受限的临床AI场景

多模态融合学习在医疗领域潜力巨大,但现有方法难以有效捕捉复杂的跨模态交互,计算成本高,且常局限于固定模态组合(如仅影像或图像与组学)。为此,本文提出轻量级可扩展的融合框架CURE,通过新颖的高效混合几何注意融合层(HyFuse)逐模态融合。HyFuse结合残差卷积模块提取多尺度特征,以低代价学习;并采用混合空间注意力机制捕捉粗粒度到细粒度的结构线索,更好保留跨模态关系。后续引入可学习的晚期融合与共享信息精炼模块,生成对模态顺序不敏感的鲁棒表示。在16个公共数据集上的广泛评估显示,CURE优于主流多模态融合方法,性能最高提升3.97%,计算开销降低达87.8%,显著提升预测效果与实用性。

原文摘要 · Abstract (English)

Multimodal fusion learning (MFL) has shown great potential in the medical domain, where we are faced with disparate data modalities such as imaging, clinical records, and omics. However, existing MFL strategies face several major challenges. First, they struggle to capture complex cross-modal interactions effectively, which in turn limits performance improvements. Second, they incur high computational costs, restricting their applicability in resource-constrained healthcare AI applications. Finally, they are often designed and evaluated for narrow, fixed modality configurations (e.g., imaging-only, or specific pairs such as image and omics), which limits evidence of their adaptability and generalizability to broader collections of heterogeneous medical modalities. To address these challenges, we propose a novel MFL framework - Cascaded Unified Representation Learning for Efficient Fusion Network (CURE) - a lightweight and scalable framework that progressively integrates various modalities through a novel efficient Hybrid Geometry Aware Fusion layer (HyFuse), where each HyFuse layer is sequentially learned for each modality, making the framework adaptable and generalizable. Within HyFuse, an efficient residual convolution module captures rich multi-scale features to ensure cost-effective learning, while a hybrid-space aware attention mixer learns coarse-to-fine structural cues to better preserve cross-modal relationships. Complementary learnable late-fusion and shared information refinement modules are then employed to learn robust modality-order-invariant shared representations, which in turn yields consistent performance improvements. Extensive evaluations on 16 public datasets show that CURE outperforms leading multimodal fusion methods, boosting performance by up to 3.97% and lowering computational costs by up to 87.8%, ensuring more effective and reliable predictions.

多模态融合医疗AI轻量化注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。