用空间填充曲线重构SAM2,提升医学图像诊断精度
Bridging the Perception-Cognition Gap:Re-engineering SAM2 with Hilbert-Mamba for Robust VLM-based Medical Diagnosis
- 将希尔伯特曲线融入Mamba架构,更好保留3D医学图像空间结构
- 在BraTS2021上分割Dice达82.35%,分类准确率78.85%
- 适合做医学影像分析的AI研究者与临床辅助系统开发者
近期研究表明,视觉语言模型(VLM)在自动化医学诊断中具有巨大潜力。然而,处理复杂的三维多模态医学图像面临挑战,特别是互补信息的有效融合以及对细微但关键病灶特征的遗漏。为此,我们提出一种新型两阶段融合框架Hilbert-VLM。该框架利用HilbertMed-SAM模块进行精确病灶分割,生成的多模态增强提示随后引导VLM完成准确疾病分类。核心创新在于对分割任意模型2(SAM2)架构的系统性重构:将希尔伯特空间填充曲线引入Mamba状态空间模型(SSM)的扫描机制,以最大化保留3D数据中的空间局部性,这对医学图像分析至关重要。同时引入新颖的希尔伯特-Mamba交叉注意力(HMCA)机制和尺度感知解码器,以捕捉细粒度细节。提示增强模块将分割掩码及其对应文本属性统一为信息密集型提示,支持VLM推理。大量实验验证了Hilbert-VLM的有效性。在BraTS2021分割基准上,其分割Dice得分为82.35%,诊断分类准确率(ACC)为78.85%。结果表明,该模型显著提升了医学VLM分析的准确性和可靠性。
原文摘要 · Abstract (English)
Recent studies suggest that Visual Language Models (VLMs) hold great potential for tasks such as automated medical diagnosis. However, processing complex three-dimensional (3D) multimodal medical images poses significant challenges - specifically, the effective integration of complementary information and the occasional oversight of subtle yet critical pathological features. To address these issues, we present a novel two-stage fusion framework termed Hilbert-VLM. This framework leverages the HilbertMed-SAM module for precise lesion segmentation, with the generated multimodal enhanced prompts then guiding the VLM toward accurate disease classification. Our key innovation lies in the systematic redesign of the Segment Anything Model 2 (SAM2) architecture: we incorporate Hilbert space-filling curves into the scanning mechanism of the Mamba State Space Model (SSM) to maximize the preservation of spatial locality in 3D data, a property critical for medical image analysis. We also introduce a novel Hilbert-Mamba Cross-Attention (HMCA) mechanism and a scale-aware decoder to capture fine-grained details. Meanwhile, the prompt enhancement module unifies segmentation masks and their corresponding textual attributes into an information-dense prompt to support VLM inference. Extensive experiments were conducted to validate the effectiveness of the Hilbert-VLM model. On the BraTS2021 segmentation benchmark, it achieves a Dice score of 82.35 percent, with a diagnostic classification accuracy (ACC) of 78.85 percent. These results demonstrate that the proposed model offers substantial potential to improve the accuracy and reliability of medical VLM-based analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。