arXiv:2411.12814cs.CV2024-11CVPR被引 47

构建3.61亿标注的医学图像分割基准数据集,支持多模态交互式分割。

Interactive Medical Image Segmentation: A Benchmark Dataset and Baseline

  • 基于视觉大模型自动生成超大规模密集标注,覆盖14种影像模态
  • 每张图平均56个分割目标,总量达3.61亿个掩码,显著超越现有数据集
  • 提供含点击、框选、文本提示的交互式分割基线模型,适配科研与临床应用

交互式医学图像分割(IMIS)长期受限于缺乏大规模、多样化且密集标注的数据集,导致模型泛化能力不足且评估不一致。本文提出IMed-361M基准数据集,是通用IMIS研究的重要进展。首先,从多个数据源收集并标准化超过640万张医学图像及其对应的真值掩码;其次,利用视觉基础模型的强大目标识别能力,自动生成每张图像的密集交互掩码,并通过严格的质量控制与粒度管理确保标注质量;不同于以往局限于特定模态或稀疏标注的数据集,IMed-361M涵盖14种模态和204个分割目标,总计3.61亿个掩码,平均每张图像56个掩码。最后,我们基于该数据集开发了支持高精度掩码生成的交互式分割基线网络,可接受点击、边界框、文本提示及其组合的交互输入。在多项医学图像分割任务上进行多角度评估,结果表明其准确性和可扩展性优于现有模型。为促进医学计算机视觉中基础模型的研究,我们已在https://github.com/uni-medical/IMIS-Bench发布IMed-361M数据集与模型。

原文摘要 · Abstract (English)

Interactive Medical Image Segmentation (IMIS) has long been constrained by the limited availability of large-scale, diverse, and densely annotated datasets, which hinders model generalization and consistent evaluation across different models. In this paper, we introduce the IMed-361M benchmark dataset, a significant advancement in general IMIS research. First, we collect and standardize over 6.4 million medical images and their corresponding ground truth masks from multiple data sources. Then, leveraging the strong object recognition capabilities of a vision foundational model, we automatically generated dense interactive masks for each image and ensured their quality through rigorous quality control and granularity management. Unlike previous datasets, which are limited by specific modalities or sparse annotations, IMed-361M spans 14 modalities and 204 segmentation targets, totaling 361 million masks-an average of 56 masks per image. Finally, we developed an IMIS baseline network on this dataset that supports high-quality mask generation through interactive inputs, including clicks, bounding boxes, text prompts, and their combinations. We evaluate its performance on medical image segmentation tasks from multiple perspectives, demonstrating superior accuracy and scalability compared to existing interactive segmentation models. To facilitate research on foundational models in medical computer vision, we release the IMed-361M and model at https://github.com/uni-medical/IMIS-Bench.

医学图像交互分割数据集大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。