arXiv:2511.19351cs.CV2025-11

构建3023张荧光显微图像数据集,推动细胞计数自动化研究。

CellFMCount: A Fluorescence Microscopy Dataset, Benchmark, and Methods for Cell Counting

  • 构建3023张免疫细胞化学图像数据集,含超43万人工标注细胞点。
  • 在10至2126个细胞/图的范围内,新方法平均误差仅22.12,优于现有最佳27.46。
  • 提供基准测试框架与SAM模型适配方案,适合生物医学图像分析研究者。

准确的细胞计数在癌症诊断、干细胞研究和免疫学等生物医学领域至关重要。人工计数耗时且易出错,促使人们采用深度学习实现自动化。然而,训练可靠的深度学习模型需要大量高质量标注数据,而人工标注成本高、耗时长,导致现有细胞计数数据集通常规模有限,常少于500张图像。本文介绍一个大规模标注数据集,包含3,023张与细胞分化相关的免疫细胞化学实验图像,涵盖超过430,000个手动标注的细胞位置。该数据集面临诸多挑战:高细胞密度、重叠与形态多样、每图细胞数量呈长尾分布,以及染色协议差异。我们在测试集(每图细胞数10至2,126)上对三类现有方法——回归型、群体计数与细胞计数技术——进行基准测试。同时评估了仅用点标注数据适配Segment Anything Model(SAM)用于显微镜细胞计数的可行性。作为案例,我们提出基于密度图的SAM改进方法(SAM-Counter),报告均方绝对误差(MAE)为22.12,优于现有方法(次优MAE为27.46)。结果表明,该数据集与基准框架对推动自动化细胞计数具有重要价值,并为未来研究提供坚实基础。

原文摘要 · Abstract (English)

Accurate cell counting is essential in various biomedical research and clinical applications, including cancer diagnosis, stem cell research, and immunology. Manual counting is labor-intensive and error-prone, motivating automation through deep learning techniques. However, training reliable deep learning models requires large amounts of high-quality annotated data, which is difficult and time-consuming to produce manually. Consequently, existing cell-counting datasets are often limited, frequently containing fewer than $500$ images. In this work, we introduce a large-scale annotated dataset comprising $3{,}023$ images from immunocytochemistry experiments related to cellular differentiation, containing over $430{,}000$ manually annotated cell locations. The dataset presents significant challenges: high cell density, overlapping and morphologically diverse cells, a long-tailed distribution of cell count per image, and variation in staining protocols. We benchmark three categories of existing methods: regression-based, crowd-counting, and cell-counting techniques on a test set with cell counts ranging from $10$ to $2{,}126$ cells per image. We also evaluate how the Segment Anything Model (SAM) can be adapted for microscopy cell counting using only dot-annotated datasets. As a case study, we implement a density-map-based adaptation of SAM (SAM-Counter) and report a mean absolute error (MAE) of $22.12$, which outperforms existing approaches (second-best MAE of $27.46$). Our results underscore the value of the dataset and the benchmarking framework for driving progress in automated cell counting and provide a robust foundation for future research and development.

细胞计数显微图像数据集SAM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。