构建了用于疟原虫检测的细粒度标注数据集,支持深度学习模型训练。
A COCO-Formatted Instance-Level Dataset for Plasmodium Falciparum Detection in Giemsa-Stained Blood Smears
- 将原始数据集升级为COCO格式,添加逐实例边界框标注
- 使用Faster R-CNN在交叉验证中实现感染细胞检测0.88的F1分数
- 适合医学图像分析、疟疾自动化诊断研究者使用
在吉姆萨染色血涂片中准确检测恶性疟原虫是可靠疟疾诊断的关键,尤其在发展中国家。基于深度学习的目标检测方法在自动疟疾诊断中展现出强大潜力,但其应用受限于缺乏详细实例级标注的数据集。本文呈现了公开可用的NIH疟疾数据集的增强版本,包含以COCO格式提供的详细边界框标注,支持目标检测训练。通过训练Faster R-CNN模型检测感染与非感染红细胞及白细胞,验证了修订标注的可靠性。在原始数据集上进行交叉验证,感染细胞检测的最高F1得分为0.88。结果表明,标注数量与一致性至关重要,且结合自动化标注优化与针对性人工修正可生成足够高质量的训练数据,实现稳健检测性能。更新后的标注数据集可通过Zenodo公开获取:https://doi.org/10.5281/zenodo.17514694。
原文摘要 · Abstract (English)
Accurate detection of Plasmodium falciparum in Giemsa-stained blood smears is an essential component of reliable malaria diagnosis, especially in developing countries. Deep learning-based object detection methods have demonstrated strong potential for automated Malaria diagnosis, but their adoption is limited by the scarcity of datasets with detailed instance-level annotations. In this work, we present an enhanced version of the publicly available NIH malaria dataset, with detailed bounding box annotations in COCO format to support object detection training. We validated the revised annotations by training a Faster R-CNN model to detect infected and non-infected red blood cells, as well as white blood cells. Cross-validation on the original dataset yielded F1 scores of up to 0.88 for infected cell detection. These results underscore the importance of annotation volume and consistency, and demonstrate that automated annotation refinement combined with targeted manual correction can produce training data of sufficient quality for robust detection performance. The updated annotations set is publicly available via Zenodo: https://doi.org/10.5281/zenodo.17514694
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。