arXiv:2508.13936eess.IVcs.CV2025-08

跨模态多器官数据训练,提升视网膜积液分割与检测精度

MMIS-Net for Retinal Fluid Segmentation and Detection

  • 设计相似性融合模块,利用标注与像素相似性选择特征
  • 在10个数据集上训练,实现积液分割平均Dice达0.83
  • 解决标签不一致问题,适合多源医疗图像任务

目的:深度学习在医学图像疾病分割与检测中表现优异,但多数方法仅基于单一来源、模态、器官或病种数据训练。大量来自不同模态、器官和疾病的公开小规模标注数据集尚未被充分利用。本文旨在挖掘这些数据的协同潜力,以提升对未见数据的表现。方法:提出新型网络MMIS-Net,包含相似性融合块,通过监督信号与像素级相似性知识选择进行特征图融合;为解决类别定义不一致与标签矛盾,构建一热标签空间处理某数据集未标注但在另一数据集中存在的类别。模型在涵盖19个器官、2种模态的10个数据集上联合训练。结果:在RETOUCH挑战赛隐藏测试集上,性能超越大型医学图像分割基础模型及其他先进算法,积液分割任务取得0.83的最高均Dice分数,绝对体积差异仅为0.035,检测任务达到完美的1.0面积曲线下积分(AUC)。结论:定量结果表明,模型因引入相似性融合块实现监督与相似性知识选择,以及使用一热标签空间缓解标签不一致问题,具备显著有效性。

原文摘要 · Abstract (English)

Purpose: Deep learning methods have shown promising results in the segmentation, and detection of diseases in medical images. However, most methods are trained and tested on data from a single source, modality, organ, or disease type, overlooking the combined potential of other available annotated data. Numerous small annotated medical image datasets from various modalities, organs, and diseases are publicly available. In this work, we aim to leverage the synergistic potential of these datasets to improve performance on unseen data. Approach: To this end, we propose a novel algorithm called MMIS-Net (MultiModal Medical Image Segmentation Network), which features Similarity Fusion blocks that utilize supervision and pixel-wise similarity knowledge selection for feature map fusion. Additionally, to address inconsistent class definitions and label contradictions, we created a one-hot label space to handle classes absent in one dataset but annotated in another. MMIS-Net was trained on 10 datasets encompassing 19 organs across 2 modalities to build a single model. Results: The algorithm was evaluated on the RETOUCH grand challenge hidden test set, outperforming large foundation models for medical image segmentation and other state-of-the-art algorithms. We achieved the best mean Dice score of 0.83 and an absolute volume difference of 0.035 for the fluids segmentation task, as well as a perfect Area Under the Curve of 1 for the fluid detection task. Conclusion: The quantitative results highlight the effectiveness of our proposed model due to the incorporation of Similarity Fusion blocks into the network's backbone for supervision and similarity knowledge selection, and the use of a one-hot label space to address label class inconsistencies and contradictions.

医学图像分割多模态视网膜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。