arXiv:2603.27460cs.CVcs.AI2026-03综述被引 3

梳理千余医学影像数据集,推动构建统一资源库以支持大模型发展

Project Imaging-X: A Survey of 1000+ Open-Access Medical Imaging Datasets for Foundation Model Development

  • 系统整理1000+开放医学影像数据集,按模态、任务等维度分类
  • 发现数据规模小、分布不均、碎片化严重,制约大模型训练
  • 提出元数据驱动融合方案,支持自动整合与交互式数据发现

基础模型在多个领域取得显著成功,主要得益于大规模、多样化和高质量数据集的支撑。然而,在医学影像领域,由于依赖临床专业知识及严格的伦理与隐私限制,构建此类数据集面临巨大挑战,导致大规模统一数据集稀缺,阻碍了强大医学基础模型的发展。本文首次系统调查超过1000个公开医学影像数据集,涵盖其模态、任务、解剖部位、标注方式、局限性及整合潜力。分析显示,现有数据集规模有限、任务高度聚焦、器官与模态分布不均,难以支撑通用且稳健的医学基础模型。为此,我们提出元数据驱动融合范式(MDFP),通过共享模态或任务整合多个小型数据孤岛,形成更大更连贯的数据资源。基于此,我们发布交互式数据发现门户,实现端到端自动化数据集成,并将所有调研数据集整理为结构化表格,清晰呈现关键特征并提供参考链接,为社区提供一个可访问、全面的资源库。本研究不仅描绘了当前医学影像数据格局,还为数据集整合提供了可操作路径,助力加速数据发现、更科学的数据构建与更强医学基础模型的开发。

原文摘要 · Abstract (English)

Foundation models have demonstrated remarkable success across diverse domains and tasks, primarily due to the thrive of large-scale, diverse, and high-quality datasets. However, in the field of medical imaging, the curation and assembling of such medical datasets are highly challenging due to the reliance on clinical expertise and strict ethical and privacy constraints, resulting in a scarcity of large-scale unified medical datasets and hindering the development of powerful medical foundation models. In this work, we present the largest survey to date of medical image datasets, covering over 1,000 open-access datasets with a systematic catalog of their modalities, tasks, anatomies, annotations, limitations, and potential for integration. Our analysis exposes a landscape that is modest in scale, fragmented across narrowly scoped tasks, and unevenly distributed across organs and modalities, which in turn limits the utility of existing medical image datasets for developing versatile and robust medical foundation models. To turn fragmentation into scale, we propose a metadata-driven fusion paradigm (MDFP) that integrates public datasets with shared modalities or tasks, thereby transforming multiple small data silos into larger, more coherent resources. Building on MDFP, we release an interactive discovery portal that enables end-to-end, automated medical image dataset integration, and compile all surveyed datasets into a unified, structured table that clearly summarizes their key characteristics and provides reference links, offering the community an accessible and comprehensive repository. By charting the current terrain and offering a principled path to dataset consolidation, our survey provides a practical roadmap for scaling medical imaging corpora, supporting faster data discovery, more principled dataset creation, and more capable medical foundation models.

医学影像数据集基础模型数据融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。