用长短记忆提升医学影像标注速度与准确率
Accelerating Volumetric Medical Image Annotation via Short-Long Memory SAM 2
- 设计长短记忆双分支结构,分离短期与长期特征存储
- 在4个数据集上平均提高0.14的Dice系数,减少60.575%纠错时间
- 适合需高效精准标注的医疗影像研究者使用
手动标注体积医学影像(如MRI、CT)耗时费力。近期视频目标分割基础模型如SAM 2可通过标注少量切片并传播掩码来加速流程,但其性能不稳定,尤其在边界区域易出现错误传播。我们提出短-长记忆SAM 2(SLM-SAM 2),引入独立的短时与长时记忆库及注意力模块,提升分割精度。在涵盖MRI、CT和超声的4个公开数据集上测试,当有5个体积用于初始适配时,平均Dice相似系数提升0.14;仅1个体积时提升0.10。该方法显著降低误传播风险,每体积纠错时间减少60.575%,为医学图像自动标注提供了更可靠方案。
原文摘要 · Abstract (English)
Manual annotation of volumetric medical images, such as magnetic resonance imaging (MRI) and computed tomography (CT), is a labor-intensive and time-consuming process. Recent advancements in foundation models for video object segmentation, such as Segment Anything Model 2 (SAM 2), offer a potential opportunity to significantly speed up the annotation process by manually annotating one or a few slices and then propagating target masks across the entire volume. However, the performance of SAM 2 in this context varies. Our experiments show that relying on a single memory bank and attention module is prone to error propagation, particularly at boundary regions where the target is present in the previous slice but absent in the current one. To address this problem, we propose Short-Long Memory SAM 2 (SLM-SAM 2), a novel architecture that integrates distinct short-term and long-term memory banks with separate attention modules to improve segmentation accuracy. We evaluate SLM-SAM 2 on four public datasets covering organs, bones, and muscles across MRI, CT, and ultrasound videos. We show that the proposed method markedly outperforms the default SAM 2, achieving an average Dice Similarity Coefficient improvement of 0.14 and 0.10 in the scenarios when 5 volumes and 1 volume are available for the initial adaptation, respectively. SLM-SAM 2 also exhibits stronger resistance to over-propagation, reducing the time required to correct propagated masks by 60.575% per volume compared to SAM 2, making a notable step toward more accurate automated annotation of medical images for segmentation model development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。