构建超大医学图文数据集Biomedica,助力通用生物医学AI发展
A Large-Scale Vision-Language Dataset Derived from Open Scientific Literature to Advance Biomedical Generalist AI
- 从600万篇文献中提取2400万图文对,含专家标注元数据
- 基于该数据集训练的AI模型性能超越现有开源系统
- 提供流式接口与搜索服务,便于科研人员快速接入使用
尽管生物医学人工智能备受关注,但高质量、多样化且大规模的数据仍是制约其发展的瓶颈。为此,我们推出了Biomedica——一个源自PubMed Central开放获取子集的开源数据集,包含超过600万篇科学论文和2400万组图像-文本配对,并附带27个元数据字段(包括专家人工标注)。为解决大规模数据访问难题,我们通过网页服务器提供可扩展的流式传输与搜索API,支持与AI系统的无缝集成。我们通过构建嵌入模型、对话式模型及检索增强型聊天代理,验证了Biomedica数据集的实用性。值得注意的是,所有自建AI模型在各自类别中均超越此前开源系统,凸显了多样、高质量、大规模生物医学数据的关键作用。
原文摘要 · Abstract (English)
Despite the excitement behind biomedical artificial intelligence (AI), access to high-quality, diverse, and large-scale data - the foundation for modern AI systems - is still a bottleneck to unlocking its full potential. To address this gap, we introduce Biomedica, an open-source dataset derived from the PubMed Central Open Access subset, containing over 6 million scientific articles and 24 million image-text pairs, along with 27 metadata fields (including expert human annotations). To overcome the challenges of accessing our large-scale dataset, we provide scalable streaming and search APIs through a web server, facilitating seamless integration with AI systems. We demonstrate the utility of the Biomedica dataset by building embedding models, chat-style models, and retrieval-augmented chat agents. Notably, all our AI models surpass previous open systems in their respective categories, underscoring the critical role of diverse, high-quality, and large-scale biomedical data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。