26万张脑部MRI数据集,助力医学影像自监督学习
A large-scale heterogeneous 3D magnetic resonance brain imaging dataset for self-supervised learning
- 整合910个公开源,构建跨设备、多模态脑部MRI数据集
- 包含55,378名受试者、77,589次扫描,覆盖广泛病理变异
- 配套预训练代码与模型,支持医疗影像自监督算法研发
我们提出FOMO260K,一个大规模异构脑部磁共振成像(MRI)数据集,包含来自910个公开来源的260,927张脑部MRI图像,涵盖77,589次扫描和55,378名受试者。数据集包含临床级与研究级图像、多种MRI序列,以及广泛的解剖与病理变异,包括存在显著脑部异常的扫描。为保留原始图像特性并降低使用门槛,仅进行最小化预处理。配套提供自监督预训练与微调代码及预训练模型。FOMO260K旨在支持医学影像中自监督学习方法的大规模开发与基准测试。
原文摘要 · Abstract (English)
We present FOMO260K, a large-scale, heterogeneous dataset of 260,927 brain Magnetic Resonance Imaging (MRI) scans from 77,589 MRI sessions and 55,378 subjects, aggregated from 910 publicly available sources. The dataset includes both clinical- and research-grade images, multiple MRI sequences, and a wide range of anatomical and pathological variability, including scans with large brain anomalies. Minimal preprocessing was applied to preserve the original image characteristics while reducing entry barriers for new users. Companion code for self-supervised pretraining and finetuning is provided, along with pretrained models. FOMO260K is intended to support the development and benchmarking of self-supervised learning methods in medical imaging at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。