用视觉Mamba提升肺肿瘤多模态分割精度
Context-Gated Cross-Modal Perception with Visual Mamba for PET-CT Lung Tumor Segmentation
- 基于视觉Mamba构建跨模态门控机制,自适应融合PET与CT特征
- 在PCLT20K数据集上达到更高分割性能,计算量更低
- 适合需要高效精准肿瘤分割的医学影像分析场景
准确的肺肿瘤分割对改善诊断和治疗规划至关重要,有效结合PET与CT的解剖与功能信息仍是重大挑战。本文提出vMambaX,一种轻量级多模态框架,通过上下文门控跨模态感知模块(CGM)融合PET与CT图像。该框架基于视觉Mamba架构,自适应增强跨模态特征交互,强调有用区域并抑制噪声。在PCLT20K数据集上的评估显示,该模型优于基线模型,同时保持更低的计算复杂度。结果表明,自适应跨模态门控在多模态肿瘤分割中有效,vMambaX展现出作为高效可扩展肺癌分析框架的潜力。代码已开源。
原文摘要 · Abstract (English)
Accurate lung tumor segmentation is vital for improving diagnosis and treatment planning, and effectively combining anatomical and functional information from PET and CT remains a major challenge. In this study, we propose vMambaX, a lightweight multimodal framework integrating PET and CT scan images through a Context-Gated Cross-Modal Perception Module (CGM). Built on the Visual Mamba architecture, vMambaX adaptively enhances inter-modality feature interaction, emphasizing informative regions while suppressing noise. Evaluated on the PCLT20K dataset, the model outperforms baseline models while maintaining lower computational complexity. These results highlight the effectiveness of adaptive cross-modal gating for multimodal tumor segmentation and demonstrate the potential of vMambaX as an efficient and scalable framework for advanced lung cancer analysis. The code is available at https://github.com/arco-group/vMambaX.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。