评估三大细胞模型在肾病理图像上的表现,探索人机协作提升分割精度的方法。
Evaluating Cell AI Foundation Models in Kidney Pathology with Human-in-the-Loop Enrichment
- 用多中心数据集测试三种前沿细胞模型的分割能力。
- 引入人机协作标注策略,显著提升模型性能且减少人工标注量。
- 发现基础模型性能不等于微调后表现,为部署提供关键参考。
训练人工智能基础模型已成为应对数字病理等现实医疗挑战的有前景的大规模学习方法。尽管许多模型已针对疾病诊断和组织量化任务,在大规模多样数据集上开发,但其在某些看似简单任务(如单一器官——肾脏中的细胞核分割)上的部署准备度仍不确定。本文通过一个精心构建的多中心、多疾病、多物种外部测试数据集,全面评估了近期细胞基础模型的表现,回答‘我们做得如何?’这一核心问题。该数据集包含2,542张肾组织全切片图像(WSIs)。选取三种先进细胞基础模型——Cellpose、StarDist和CellViT进行评估。同时,为回答‘我们如何改进?’,提出并测试了基于人类反馈的人机协同数据增强策略,旨在以最小化像素级标注代价提升模型性能。实验结果表明,所有三种模型在使用增强数据微调后均优于基线。有趣的是,基线F1分数最高的模型在微调后并未取得最佳分割效果。本研究为面向真实数据应用的细胞视觉基础模型开发与部署建立了基准。
原文摘要 · Abstract (English)
Training AI foundation models has emerged as a promising large-scale learning approach for addressing real-world healthcare challenges, including digital pathology. While many of these models have been developed for tasks like disease diagnosis and tissue quantification using extensive and diverse training datasets, their readiness for deployment on some arguably simplest tasks, such as nuclei segmentation within a single organ (e.g., the kidney), remains uncertain. This paper seeks to answer this key question, "How good are we?", by thoroughly evaluating the performance of recent cell foundation models on a curated multi-center, multi-disease, and multi-species external testing dataset. Additionally, we tackle a more challenging question, "How can we improve?", by developing and assessing human-in-the-loop data enrichment strategies aimed at enhancing model performance while minimizing the reliance on pixel-level human annotation. To address the first question, we curated a multicenter, multidisease, and multispecies dataset consisting of 2,542 kidney whole slide images (WSIs). Three state-of-the-art (SOTA) cell foundation models-Cellpose, StarDist, and CellViT-were selected for evaluation. To tackle the second question, we explored data enrichment algorithms by distilling predictions from the different foundation models with a human-in-the-loop framework, aiming to further enhance foundation model performance with minimal human efforts. Our experimental results showed that all three foundation models improved over their baselines with model fine-tuning with enriched data. Interestingly, the baseline model with the highest F1 score does not yield the best segmentation outcomes after fine-tuning. This study establishes a benchmark for the development and deployment of cell vision foundation models tailored for real-world data applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。