arXiv:2512.15921eess.IVcs.CV2025-12

无标注数据时,用统一标准评估多个医学图像分割模型表现。

In search of truth: Evaluating concordance of AI-based anatomy segmentation models

  • 将不同模型输出统一为标准格式,实现结构化对比
  • 在NLST数据集上验证,肺部分割一致性强,脊椎肋骨则差异大
  • 提供可视化工具和脚本,适合医学影像研究者使用

基于人工智能的解剖结构分割方法可自动化处理大规模影像数据。随着功能相似模型日益增多,缺乏真实标注数据时的评估成为挑战。本文提出一种实用框架:将分割结果统一为标准、可互操作的表示形式,支持基于术语的一致性标注。通过扩展3D Slicer,实现分割结果的便捷加载与对比,并结合交互式汇总图和OHIF Viewer浏览器可视化展示。以公开的国家肺癌筛查试验(NLST)CT数据集为例,对6个开源模型(TotalSegmentator 1.5/2.6、Auto3DSeg、MOOSE、MultiTalent、CADS)在31个解剖结构(肺、脊椎、肋骨、心脏等)上的分割性能进行评估。结果表明该框架可自动化完成加载、逐结构检查与跨模型比较;初步结果显示其能快速发现并审查异常结果。对比显示部分结构(如肺)分割一致性高,但部分模型对脊椎或肋骨产生无效分割。相关资源已发布于https://imagingdatacommons.github.io/segmentation-comparison/,包含分割统一脚本、汇总图与可视化工具,助力无真实标注情况下的模型评估与选型。

原文摘要 · Abstract (English)

Purpose AI-based methods for anatomy segmentation can help automate characterization of large imaging datasets. The growing number of similar in functionality models raises the challenge of evaluating them on datasets that do not contain ground truth annotations. We introduce a practical framework to assist in this task. Approach We harmonize the segmentation results into a standard, interoperable representation, which enables consistent, terminology-based labeling of the structures. We extend 3D Slicer to streamline loading and comparison of these harmonized segmentations, and demonstrate how standard representation simplifies review of the results using interactive summary plots and browser-based visualization using OHIF Viewer. To demonstrate the utility of the approach we apply it to evaluating segmentation of 31 anatomical structures (lungs, vertebrae, ribs, and heart) by six open-source models - TotalSegmentator 1.5 and 2.6, Auto3DSeg, MOOSE, MultiTalent, and CADS - for a sample of Computed Tomography (CT) scans from the publicly available National Lung Screening Trial (NLST) dataset. Results We demonstrate the utility of the framework in enabling automating loading, structure-wise inspection and comparison across models. Preliminary results ascertain practical utility of the approach in allowing quick detection and review of problematic results. The comparison shows excellent agreement segmenting some (e.g., lung) but not all structures (e.g., some models produce invalid vertebrae or rib segmentations). Conclusions The resources developed are linked from https://imagingdatacommons.github.io/segmentation-comparison/ including segmentation harmonization scripts, summary plots, and visualization tools. This work assists in model evaluation in absence of ground truth, ultimately enabling informed model selection.

医学图像分割评估AI模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。