arXiv:2501.07171cs.CVcs.CL2025-01CVPR被引 50

构建首个覆盖全领域生物医学图像文本对的开源数据集,助力通用医学视觉语言模型发展

BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature

论文配图:BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
图 1 · 摘自论文原文
  • 从600万篇文献中提取2400万组图像-文本对,自动标注并结构化
  • 基于该数据集训练的模型在40项任务上达顶尖性能,零样本分类平均提升6.56%
  • 提供免下载27TB数据的流式预训练模型,适合医疗AI研究者快速部署

视觉语言模型的发展依赖大规模多样化的多模态数据集。然而,通用生物医学视觉语言模型的进步受限于缺乏跨生物学与医学领域的注释公开数据集。现有工作局限于狭窄领域,未能涵盖科学文献中的全部生物医学知识多样性。为此,我们提出BIOMEDICA,一个可扩展的开源框架,用于从PubMed Central开放获取子集提取、注释并序列化全部内容,生成易于使用的公开数据集。该框架产出包含超过2400万唯一图像-文本对的综合性档案,源自超过600万篇文章,附带元数据和专家引导标注。通过持续流式预训练释放BMCA-CLIP系列CLIP风格模型,无需本地下载27TB数据。平均而言,我们的模型在40项任务(涵盖病理学、放射学、眼科、皮肤病学、外科、分子生物学、寄生虫学和细胞生物学)上达到最先进水平,在零样本分类中平均提升6.56%(皮肤病学最高达29.8%,眼科达17.5%),图像-文本检索能力更强,且计算量减少10倍。为促进可复现性与协作,我们向研究社区开放代码库与数据集。

原文摘要 · Abstract (English)

The development of vision-language models (VLMs) is driven by large-scale and diverse multimodal datasets. However, progress toward generalist biomedical VLMs is limited by the lack of annotated, publicly accessible datasets across biology and medicine. Existing efforts are restricted to narrow domains, missing the full diversity of biomedical knowledge encoded in scientific literature. To address this gap, we introduce BIOMEDICA, a scalable, open-source framework to extract, annotate, and serialize the entirety of the PubMed Central Open Access subset into an easy-to-use, publicly accessible dataset. Our framework produces a comprehensive archive with over 24 million unique image-text pairs from over 6 million articles. Metadata and expert-guided annotations are also provided. We demonstrate the utility and accessibility of our resource by releasing BMCA-CLIP, a suite of CLIP-style models continuously pre-trained on the BIOMEDICA dataset via streaming, eliminating the need to download 27 TB of data locally. On average, our models achieve state-of-the-art performance across 40 tasks - spanning pathology, radiology, ophthalmology, dermatology, surgery, molecular biology, parasitology, and cell biology - excelling in zero-shot classification with a 6.56% average improvement (as high as 29.8% and 17.5% in dermatology and ophthalmology, respectively), and stronger image-text retrieval, all while using 10x less compute. To foster reproducibility and collaboration, we release our codebase and dataset for the broader research community.

生物医学图像文本对视觉语言模型开源数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。