通过自适应选帧引擎提升多模态医学图像分割的准确性与效率
Adaptive Interactive Segmentation for Multimodal Medical Imaging via Selection Engine
- 基于SAM2构建策略驱动模型,引入动态选帧机制优化提示选择
- 在10个数据集、7种模态上实现稳定高精度分割,泛化能力强
- 无需专业知识即可自动选帧,适合临床实时交互场景
在医学图像分析中,快速、高效且准确的分割对自动化诊断与治疗至关重要。尽管深度学习显著提升了分割精度,现有模型在处理多模态医学影像时仍面临适应性差、泛化能力弱的问题,主要源于成像模态间差异大及数据本身复杂。为此,我们提出基于SAM2的策略驱动交互分割模型(SISeg),通过集成自适应帧选择引擎(AFSE)提升跨模态分割性能。AFSE在2D图像序列推理中自动选择最优提示帧,缓解内存瓶颈,无需额外医学知识,同时通过交互反馈机制增强模型可解释性。我们在涵盖7种代表性医学成像模态的10个数据集上进行了广泛实验,验证了SISeg在多模态任务中的强适应性与泛化能力。项目主页与代码将公开。
原文摘要 · Abstract (English)
In medical image analysis, achieving fast, efficient, and accurate segmentation is essential for automated diagnosis and treatment. Although recent advancements in deep learning have significantly improved segmentation accuracy, current models often face challenges in adaptability and generalization, particularly when processing multi-modal medical imaging data. These limitations stem from the substantial variations between imaging modalities and the inherent complexity of medical data. To address these challenges, we propose the Strategy-driven Interactive Segmentation Model (SISeg), built on SAM2, which enhances segmentation performance across various medical imaging modalities by integrating a selection engine. To mitigate memory bottlenecks and optimize prompt frame selection during the inference of 2D image sequences, we developed an automated system, the Adaptive Frame Selection Engine (AFSE). This system dynamically selects the optimal prompt frames without requiring extensive prior medical knowledge and enhances the interpretability of the model's inference process through an interactive feedback mechanism. We conducted extensive experiments on 10 datasets covering 7 representative medical imaging modalities, demonstrating the SISeg model's robust adaptability and generalization in multi-modal tasks. The project page and code will be available at: [URL].
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。