系统分析54个公开脑部MRI数据集,揭示数据差异对模型训练的影响。
A Structured Review and Quantitative Profiling of Public Brain MRI Datasets for Foundation Model Development
- 梳理54个数据集的模态、疾病覆盖与规模,发现健康人群数据远多于临床数据。
- 量化15个数据集图像的体素间距与强度分布,显示显著异质性。
- 验证预处理无法完全消除跨数据集偏差,需设计适应性模型策略。
脑部MRI基础模型的发展依赖于数据的规模、多样性与一致性,但现有系统评估仍不足。本研究分析了54个公开脑部MRI数据集,涵盖超过538,031例扫描,从数据集层面分析模态组成、疾病覆盖与规模,发现大型健康队列与小型临床群体间存在明显失衡;在图像层面,对15个代表性数据集的体素间距、方向与强度分布进行量化,揭示显著异质性,可能影响表征学习;进一步评估预处理流程(如强度归一化、偏场校正、头骨剥离、空间配准与插值)对体素统计与几何结构的影响,虽能提升数据内一致性,但跨数据集残余差异依然存在;通过3D DenseNet121的特征空间案例研究,证实标准化预处理后仍存在可测量的协变量偏移,表明仅靠数据调和无法消除跨数据集偏差。研究为公共脑部MRI资源提供了统一的变异表征,强调在构建通用脑部MRI基础模型时应采用预处理感知与领域自适应策略。
原文摘要 · Abstract (English)
The development of foundation models for brain MRI depends critically on the scale, diversity, and consistency of available data, yet systematic assessments of these factors remain scarce. In this study, we analyze 54 publicly accessible brain MRI datasets encompassing over 538,031 to provide a structured, multi-level overview tailored to foundation model development. At the dataset level, we characterize modality composition, disease coverage, and dataset scale, revealing strong imbalances between large healthy cohorts and smaller clinical populations. At the image level, we quantify voxel spacing, orientation, and intensity distributions across 15 representative datasets, demonstrating substantial heterogeneity that can influence representation learning. We then perform a quantitative evaluation of preprocessing variability, examining how intensity normalization, bias field correction, skull stripping, spatial registration, and interpolation alter voxel statistics and geometry. While these steps improve within-dataset consistency, residual differences persist between datasets. Finally, feature-space case study using a 3D DenseNet121 shows measurable residual covariate shift after standardized preprocessing, confirming that harmonization alone cannot eliminate inter-dataset bias. Together, these analyses provide a unified characterization of variability in public brain MRI resources and emphasize the need for preprocessing-aware and domain-adaptive strategies in the design of generalizable brain MRI foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。