arXiv:2608.19788eess.IVcs.CV2026-08

首个支持异构缺失模态的联邦弱监督肿瘤分割框架,实现高精度医学图像分割。

MOSAIC: Modality-agnostic Spectral Alignment for Federated Image-level Weakly Supervised Tumor Segmentation under Client-specific Missing Modalities

论文配图:MOSAIC: Modality-agnostic Spectral Alignment for Federated Image-level Weakly Supervised Tumor Segmentation under Client-specific Missing Modalities
图 1 · 摘自论文原文
  • 通过无模态身份感知的对齐模块融合不同机构的不完整多模态数据
  • 利用频域统计量解决跨机构分布偏移,在FeTS2022上达0.84 Dice
  • 适用于隐私受限场景,新机构加入无需重训练,仅需0.01-0.04增量

临床中可信的多模态融合需应对机构间不完整且异构的模态子集问题,而隐私限制又禁止数据集中共享。联邦学习(FL)缓解了数据共享约束,但面临客户端特定缺失模态的问题,即各机构仅有部分多模态数据,导致融合质量下降和分割性能降低。尽管已有研究分别探讨了联邦学习与弱监督,但在异构缺失模态下结合图像级标签的联合方法尚未被探索。本文提出首个无模态依赖的联邦弱监督二值肿瘤分割框架MOSAIC:引入客户端特定的模态对齐模块,将可用通道融合至共享潜在空间而不依赖模态身份;设计谱原型对齐损失,利用紧凑非可逆频域统计量调和跨客户端分布偏移;构建专用联邦精炼网络,将类激活图伪标签去噪为准确掩码,突破弱监督精度瓶颈。在三个多中心脑肿瘤基准(FeTS2022、BraTS-MEN、BraTS-SSA)上的实验表明,相比所有图像、框、点监督基线均有显著提升,仅用图像级标签即接近全监督性能,在FeTS2022上达到0.84 Dice;动态新增客户端可在不重训练情况下实现0.01–0.04 Dice增量。代码已开源。

原文摘要 · Abstract (English)

Trustworthy multimodal fusion in clinical settings requires handling incomplete and heterogeneous modality subsets across institutions, where privacy constraints prohibit centralized data sharing. Federated learning (FL) mitigates data-sharing constraints but suffers from client-specific missing modalities, where institutions possess incomplete multimodal subsets, degrading fusion quality and segmentation performance. While FL and weak supervision have been studied separately, their joint use with image-level labels under heterogeneous missing modalities remains unaddressed. We propose \textbf{MOSAIC}, the first modality-agnostic federated framework for weakly supervised binary tumor segmentation under client-specific missing modalities. We introduce a client-specific modality-alignment module that fuses available channels into a shared latent space without prior knowledge of modality identity, a spectral prototype alignment loss that reconciles cross-client distribution shift using compact non-invertible frequency-domain statistics, and a dedicated federated refinement network that denoises the resulting CAM pseudo-labels into accurate masks, breaking the accuracy ceiling of weak supervision. Experiments on three multi-institutional brain tumor benchmarks (FeTS2022, BraTS-MEN, and BraTS-SSA) demonstrate significant improvements over all image, box, and point-supervised baselines, approaching fully supervised accuracy using only image-level labels and reaching 0.84 Dice on FeTS2022. Dynamic new client addition enables previously unseen institutions to join an already-trained federation within 0.01-0.04 Dice without retraining. Code is available at https://github.com/Tarun2201/MOSAIC.

联邦学习弱监督肿瘤分割多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。