arXiv:2512.04564cs.CV2025-12综述

如何构建高质量显微图像数据集,助力病理深度学习模型可靠落地。

Dataset creation for supervised deep learning-based analysis of microscopic images -- review of important considerations and recommendations

  • 系统梳理图像采集、标注工具选型与标注生成的关键步骤。
  • 强调标注质量需满足正确性、完整性、一致性,避免偏差影响模型性能。
  • 提供标准化操作流程,适合科研人员与医疗AI团队参考实践。

监督式深度学习在显微图像自动分析中备受关注,其模型的开发与验证高度依赖高质量、大规模数据集。然而,数据集构建过程复杂且耗时,常受制于时间限制、领域差异及标注环节的偏倚风险。本文全面回顾数据集创建的关键步骤:1)图像采集,2)标注软件选择,3)标注生成。除确保足够数量的图像外,还需关注由切片制备与数字化带来的图像变异性(领域偏移),若训练数据未充分覆盖,将导致算法错误。标注质量应遵循“三C”标准:正确性、完整性、一致性。文章探讨通过先进方法提升标注质量,缓解单个标注者局限。附录提供标准操作流程(SOP),指导数据集开发。同时强调开放数据集对推动创新与研究可复现性的关键作用。通过解决挑战并提出实用建议,本综述旨在促进高质量、大规模数据集的创建与共享,最终支持通用性强、鲁棒性高的病理深度学习模型发展。

原文摘要 · Abstract (English)

Supervised deep learning (DL) receives great interest for automated analysis of microscopic images with an increasing body of literature supporting its potential. The development and validation of those DL models relies heavily on the availability of high-quality, large-scale datasets. However, creating such datasets is a complex and resource-intensive process, often hindered by challenges such as time constraints, domain variability, and risks of bias in image collection and label creation. This review provides a comprehensive guide to the critical steps in dataset creation, including: 1) image acquisition, 2) selection of annotation software, and 3) annotation creation. In addition to ensuring a sufficiently large number of images, it is crucial to address sources of image variability (domain shifts) - such as those related to slide preparation and digitization - that could lead to algorithmic errors if not adequately represented in the training data. Key quality criteria for annotations are the three "C"s: correctness, completeness, and consistency. This review explores methods to enhance annotation quality through the use of advanced techniques that mitigate the limitations of single annotators. To support dataset creators, a standard operating procedure (SOP) is provided as supplemental material, outlining best practices for dataset development. Furthermore, the article underscores the importance of open datasets in driving innovation and enhancing reproducibility of DL research. By addressing the challenges and offering practical recommendations, this review aims to advance the creation of and availability to high-quality, large-scale datasets, ultimately contributing to the development of generalizable and robust DL models for pathology applications.

数据集构建病理图像深度学习标注质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。