构建了用于结直肠疾病生物标志物发现的高粒度组织类型标注数据集
ADPv2: A Hierarchical Histological Tissue Type-Annotated Dataset for Potential Biomarker Discovery of Colorectal Disease
- 基于32类分层组织类型,标注2万张结肠活检图像
- 采用VMamba模型实现0.88的多标签分类mAP,性能优异
- 支持结肠癌病理机制分析,适合临床与算法研究者使用
计算病理学利用组织病理图像提升临床诊断的精确性与可重复性。然而,因需专业知识和高昂标注成本,现有公开数据集中能提供细粒度组织类型(HTT)层级标注的仍十分稀缺。已有数据集如数字病理学图谱(ADP)虽涵盖多器官的丰富组织类型标注,但难以支持特定器官疾病的深入研究。为此,我们推出了聚焦胃肠道病理学的ADPv2数据集,包含20,004张来自健康结肠活检切片的图像块,依据三级结构的32类组织类型进行标注。此外,我们在该数据集上训练了一个两阶段的多标签表征学习模型,采用VMamba架构,实现0.88的平均精度均值(mAP)。最后,通过分析模型在不同结肠疾病组织上的预测行为,揭示了结肠癌发展的两种病理路径的统计模式,验证了本数据集在器官特异性深度研究与潜在生物标志物发现中的能力。数据集已公开于https://zenodo.org/records/15307021。
原文摘要 · Abstract (English)
Computational pathology (CoPath) leverages histopathology images to enhance diagnostic precision and reproducibility in clinical pathology. However, publicly available datasets for CoPath that are annotated with extensive histological tissue type (HTT) taxonomies at a granular level remain scarce due to the significant expertise and high annotation costs required. Existing datasets, such as the Atlas of Digital Pathology (ADP), address this by offering diverse HTT annotations generalized to multiple organs, but limit the capability for in-depth studies on specific organ diseases. Building upon this foundation, we introduce ADPv2, a novel dataset focused on gastrointestinal histopathology. Our dataset comprises 20,004 image patches derived from healthy colon biopsy slides, annotated according to a hierarchical taxonomy of 32 distinct HTTs of 3 levels. Furthermore, we train a multilabel representation learning model following a two-stage training procedure on our ADPv2 dataset. We leverage the VMamba architecture and achieving a mean average precision (mAP) of 0.88 in multilabel classification of colon HTTs. Finally, we show that our dataset is capable of an organ-specific in-depth study for potential biomarker discovery by analyzing the model's prediction behavior on tissues affected by different colon diseases, which reveals statistical patterns that confirm the two pathological pathways of colon cancer development. Our dataset is publicly available at https://zenodo.org/records/15307021
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。