构建首个多器官病理图像数据集,助力精准医疗AI发展
SPIDER: A Comprehensive Multi-Organ Supervised Pathology Dataset and Baseline Models
- 构建覆盖皮肤、结肠、胸腔和乳腺的多器官病理图像数据集
- 引入专家标注与上下文补丁,分类准确率显著提升
- 提供可复现的基准模型,适合病理AI研究者使用
推动计算病理学中的AI发展需要大规模、高质量且多样化的数据集,但现有公开数据集在器官多样性、类别覆盖或标注质量方面仍有限。为此,我们推出了SPIDER(Supervised Pathology Image-DEscription Repository),这是目前最大的公开补丁级病理图像数据集,涵盖皮肤、结直肠、胸腔和乳腺四种器官,每类均有全面的病理类别覆盖。SPIDER提供由专家病理医生验证的高质量标注,并包含周围上下文补丁,通过空间上下文信息提升分类性能。同时,我们基于Hibou-L基础模型作为特征提取器,结合注意力分类头,训练了基准模型,在多个组织类别上达到领先水平,为未来数字病理研究提供了强有力基准。该模型还可实现关键区域快速定位、组织定量分析,并为多模态方法奠定基础。数据集与训练模型均已开源,旨在促进研究可复现性与病理AI发展。访问地址:https://github.com/HistAI/SPIDER
原文摘要 · Abstract (English)
Advancing AI in computational pathology requires large, high-quality, and diverse datasets, yet existing public datasets are often limited in organ diversity, class coverage, or annotation quality. To bridge this gap, we introduce SPIDER (Supervised Pathology Image-DEscription Repository), the largest publicly available patch-level dataset covering multiple organ types, including Skin, Colorectal, Thorax, and Breast with comprehensive class coverage for each organ. SPIDER provides high-quality annotations verified by expert pathologists and includes surrounding context patches, which enhance classification performance by providing spatial context. Alongside the dataset, we present baseline models trained on SPIDER using the Hibou-L foundation model as a feature extractor combined with an attention-based classification head. The models achieve state-of-the-art performance across multiple tissue categories and serve as strong benchmarks for future digital pathology research. Beyond patch classification, the model enables rapid identification of significant areas, quantitative tissue metrics, and establishes a foundation for multimodal approaches. Both the dataset and trained models are publicly available to advance research, reproducibility, and AI-driven pathology development. Access them at: https://github.com/HistAI/SPIDER
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。