arXiv:2601.14406cs.CVeess.IV2026-01中稿 · ed被引 1

用视觉语言模型自动评估142个解剖结构的标注质量,提升医学分割数据可信度。

Large-Scale Label Quality Assessment for Medical Segmentation via a Vision-Language Judge and Synthetic Data

  • 基于400万图像-标签对训练轻量级视觉语言模型,实现快速质量评估
  • 与真实骰子相似系数相关性达0.902,3D掩码评估仅需0.06秒
  • 可显著降低标注成本与质检时间,适用于大规模医疗数据清洗

大规模医学分割数据集常混合人工标注与伪标签,质量参差不齐,影响模型训练与评估效果。为解决此问题,本文提出SegAE(Segmentation Assessment Engine),一种轻量级视觉语言模型(VLM),可自动预测142个解剖结构的标注质量。该模型在超过四百万张图像-标签对及其质量评分上进行训练,与真实骰子相似系数的相关系数达0.902,单个3D掩码评估耗时仅0.06秒。分析显示公共数据集中普遍存在低质量标注;使用SegAE可提升主动学习与半监督学习中的数据效率,使数据标注成本降低三分之一,每标签质检时间减少70%。该工具为大规模医学分割数据集的质量控制提供了简单有效的解决方案。数据集、模型权重及代码已开源:https://github.com/Schuture/SegAE。

原文摘要 · Abstract (English)

Large-scale medical segmentation datasets often combine manual and pseudo-labels of uneven quality, which can compromise training and evaluation. Low-quality labels may hamper performance and make the model training less robust. To address this issue, we propose SegAE (Segmentation Assessment Engine), a lightweight vision-language model (VLM) that automatically predicts label quality across 142 anatomical structures. Trained on over four million image-label pairs with quality scores, SegAE achieves a high correlation coefficient of 0.902 with ground-truth Dice similarity and evaluates a 3D mask in 0.06s. SegAE shows several practical benefits: (I) Our analysis reveals widespread low-quality labeling across public datasets; (II) SegAE improves data efficiency and training performance in active and semi-supervised learning, reducing dataset annotation cost by one-third and quality-checking time by 70% per label. This tool provides a simple and effective solution for quality control in large-scale medical segmentation datasets. The dataset, model weights, and codes are released at https://github.com/Schuture/SegAE.

医学图像标注质量视觉语言模型数据清洗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。