为病理图像生成标准化组织图谱,助力AI高效构建高质量数据集。
From slides to AI-ready maps: Standardized multi-layer tissue maps as metadata for artificial intelligence in digital pathology
- 构建三层数字图谱:来源、组织类型、病理变化,统一描述切片内容
- 通过AI自动提取信息,实现百万级切片的快速分类与检索
- 适合医学影像研究者、病理AI开发者及大规模数据集构建团队
全切片图像(WSI)是通过扫描载玻片上生物样本(如组织切片或细胞样本)在多倍率下生成的高分辨率数字图像,广泛用于人工智能算法开发。尽管在病理诊断和癌症研究中至关重要,其应用还覆盖神经学、兽医、血液学、微生物学、皮肤病学、药理学、毒理学、免疫学及法医学等领域。当前缺乏统一的图像元数据标准,数据筛选主要依赖人工检查,难以应对包含数百万对象的大规模集合。本文提出一种通用框架,生成基于通用语法与语义的二维索引图(组织图谱),实现不同数据库间的互操作性。组织图谱分为三层:来源、组织类型、病理改变,每层将切片区域归类,提供可直接用于AI的元数据。我们通过基于AI的元数据提取方法生成组织图谱,并将其整合至WSI档案库中,显著提升档案搜索能力,加速高质量、平衡且目标明确的数据集构建,推动人工智能训练、验证及癌症研究进程。
原文摘要 · Abstract (English)
A Whole Slide Image (WSI) is a high-resolution digital image created by scanning an entire glass slide containing a biological specimen, such as tissue sections or cell samples, at multiple magnifications. These images are digitally viewable, analyzable, and shareable, and are widely used for Artificial Intelligence (AI) algorithm development. WSIs play an important role in pathology for disease diagnosis and oncology for cancer research, but are also applied in neurology, veterinary medicine, hematology, microbiology, dermatology, pharmacology, toxicology, immunology, and forensic science. When assembling cohorts for AI training or validation, it is essential to know the content of a WSI. However, no standard currently exists for this metadata, and such a selection has largely relied on manual inspection, which is not suitable for large collections with millions of objects. We propose a general framework to generate 2D index maps (tissue maps) that describe the morphological content of WSIs using common syntax and semantics to achieve interoperability between catalogs. The tissue maps are structured in three layers: source, tissue type, and pathological alterations. Each layer assigns WSI segments to specific classes, providing AI-ready metadata. We demonstrate the advantages of this standard by applying AI-based metadata extraction from WSIs to generate tissue maps and integrating them into a WSI archive. This integration enhances search capabilities within WSI archives, thereby facilitating the accelerated assembly of high-quality, balanced, and more targeted datasets for AI training, validation, and cancer research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。