arXiv:2508.16882eess.IVcs.CV2025-08

通过多模态图像融合提升喉部肿瘤分割精度

Multimodal Medical Endoscopic Image Analysis via Progressive Disentangle-aware Contrastive Learning

  • 设计渐进解耦对比学习框架,分离不同模态特征
  • 在多个数据集上实现优于现有方法的分割准确率
  • 适合临床辅助诊断与多模态医学图像分析研究者

准确分割喉咽部肿瘤对精准诊断和治疗规划至关重要。传统单模态成像方法难以捕捉肿瘤复杂的解剖与病理特征。本文提出基于'对齐-解耦-融合'机制的多模态表示学习框架,无缝整合2D白光成像(WLI)与窄带成像(NBI)图像对,提升分割性能。核心在于多尺度分布对齐,通过跨多个Transformer层的特征对齐缓解模态差异。同时,采用渐进式特征解耦策略,结合初步解耦与解耦感知对比学习,有效分离模态特异性和共享特征,实现鲁棒的多模态对比学习与高效语义融合。在多个数据集上的全面实验表明,本方法持续优于现有先进方法,在多种真实临床场景中均表现出更优的准确性。

原文摘要 · Abstract (English)

Accurate segmentation of laryngo-pharyngeal tumors is crucial for precise diagnosis and effective treatment planning. However, traditional single-modality imaging methods often fall short of capturing the complex anatomical and pathological features of these tumors. In this study, we present an innovative multi-modality representation learning framework based on the `Align-Disentangle-Fusion' mechanism that seamlessly integrates 2D White Light Imaging (WLI) and Narrow Band Imaging (NBI) pairs to enhance segmentation performance. A cornerstone of our approach is multi-scale distribution alignment, which mitigates modality discrepancies by aligning features across multiple transformer layers. Furthermore, a progressive feature disentanglement strategy is developed with the designed preliminary disentanglement and disentangle-aware contrastive learning to effectively separate modality-specific and shared features, enabling robust multimodal contrastive learning and efficient semantic fusion. Comprehensive experiments on multiple datasets demonstrate that our method consistently outperforms state-of-the-art approaches, achieving superior accuracy across diverse real clinical scenarios.

医学图像多模态学习分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。