arXiv:2505.12120eess.IVcs.CV2025-05被引 16

构建6万张病理切片数据集,推动精准医疗AI发展

HISTAI: An Open-Source, Large-Scale Whole Slide Image Dataset for Computational Pathology

  • 整合6万余张多组织类型病理切片,支持多模态分析
  • 每例均含诊断、人口统计、病理标注及标准编码等完整临床信息
  • 开源开放,助力可复现的临床级病理AI研究

数字病理学中人工智能与基础模型的发展凸显了大规模、多样化且标注丰富的数据集的重要性。然而,现有公开的全切片图像(WSI)数据集普遍规模不足、组织类型单一、临床元数据不全,制约了AI模型的鲁棒性与泛化能力。为此,我们推出HISTAI数据集,一个大型、多模态、开源的全切片图像集合,包含超过60,000张来自多种组织类型的切片。每个病例均配有详尽的临床元数据,包括诊断结果、人口统计信息、详细病理标注及标准化诊断编码。该数据集旨在填补现有资源的空白,推动创新、可复现性及临床相关的计算病理学解决方案发展。数据集可通过 https://github.com/HistAI/HISTAI 获取。

原文摘要 · Abstract (English)

Recent advancements in Digital Pathology (DP), particularly through artificial intelligence and Foundation Models, have underscored the importance of large-scale, diverse, and richly annotated datasets. Despite their critical role, publicly available Whole Slide Image (WSI) datasets often lack sufficient scale, tissue diversity, and comprehensive clinical metadata, limiting the robustness and generalizability of AI models. In response, we introduce the HISTAI dataset, a large, multimodal, open-access WSI collection comprising over 60,000 slides from various tissue types. Each case in the HISTAI dataset is accompanied by extensive clinical metadata, including diagnosis, demographic information, detailed pathological annotations, and standardized diagnostic coding. The dataset aims to fill gaps identified in existing resources, promoting innovation, reproducibility, and the development of clinically relevant computational pathology solutions. The dataset can be accessed at https://github.com/HistAI/HISTAI.

病理图像医学AI开源数据集数字病理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。