arXiv:2605.21861cs.CVcs.AI2026-05KDD

提出模块化架构DEX,解决多模态医学图像模型的特征偏差问题。

Learning Emergent Modular Representations in Multi-modality Medical Vision Foundation Models

论文配图:Learning Emergent Modular Representations in Multi-modality Medical Vision Foundation Models
图 1 · 摘自论文原文
  • 采用专家-导师机制,动态分配不同模态的特征学习任务。
  • 在26个下游任务中表现更优,优化更稳定,迁移能力更强。
  • 适合构建通用多模态医学AI系统的研究者和开发者。

多模态医学视觉基础模型面临不同成像模态间显著非独立同分布的特征统计差异问题。对异构数据进行单一自监督优化会引发冲突梯度,导致表示退化为模态主导的捷径。本文将此失败重构为专业化与协同性之间平衡失衡的涌现模块性问题,提出专家-导师(DEX)模块化网络,通过堆叠模块显式调控这一动态。每个模块包含一组由图像级激活策略动态调整的专家,自主适应模态主导的统计特性;同时配备由组指数移动平均更新的导师,将多专家知识提炼至共享空间,实现跨模态语义整合,推动模块化表征的涌现。我们构建了新的基准Medical Vision Universe,涵盖400万张图像、10种成像模态,为当前覆盖最广的医学视觉基础模型预训练提供支持。在26个下游任务上的大量评估表明,DEX具备更优的优化行为和迁移性能,标志着向通用多模态医学人工智能迈出原则性一步。代码与数据集将公开于https://github.com/YutingHe-list/DEX。

原文摘要 · Abstract (English)

Multi-modality medical vision (MV) foundation models (FM) are fundamentally challenged by pronounced Non-IID feature statistics across heterogeneous imaging modalities. Monolithic self-supervised optimization on such data induces conflicting gradients, driving representations to collapse toward modality-dominant shortcuts. This work reframes this failure as an imbalance between specialization and coordination in emergent modularity, and proposes Director-Experts (DEX), a modular network that explicitly regulates these dynamics in stacked modules. Each DEX module comprises a pool of experts, dynamically adapted by our image-wise activation strategy, autonomously specializing in modality-dominant statistics, together with a director, updated via our group exponential moving average, which distills multi-expert knowledge into a shared space for semantic integration across modalities, thus driving the emergence of modular representations. We curate a new benchmark, Medical Vision Universe, over 4 million images across 10 modalities, which provides a FM-level pre-training with the broadest coverage of distinct imaging modalities to our DEX. Extensive evaluations on 26 downstream tasks demonstrate improved optimization behavior and transferability, indicating DEX as a principled step toward general-purpose multi-modality medical AI. Our code and dataset will be opened at https://github.com/YutingHe-list/DEX.

多模态医学影像模块化基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。