构建首个眼科视觉语言数据集,助力医学图像理解模型训练
PubMed-Ophtha: An open resource for training ophthalmology vision-language models on scientific literature

- 从1.5万篇开放论文中提取超十万张高清眼科图像与对应图文对
- 图像按图面板拆分并标注成像模态和标记信息,准确率达90%以上
- 开源全部标注数据、模型和生成流程,支持可复现研究
视觉语言模型在眼科领域潜力巨大,但高质量的大规模图像-文本数据集仍稀缺。本文提出PubMed-Ophtha,一个包含102,023对眼科图像-标题的分层数据集,源自15,842篇来自PubMed Central的开放获取文章。不同于现有数据集,本研究直接从文章PDF中提取全分辨率图像,并将其分解为独立图面板、面板标识符及单个图像。每张图像均标注成像模态(如彩色眼底摄影、光学相干断层扫描等)及是否存在注释标记(如箭头)。通过两阶段LLM方法将图注切分为面板级子标题,在人工标注数据上达到平均句子BLEU得分0.913。面板与图像检测模型的[email protected]分别为0.909和0.892,图块提取的中位IoU达0.997。为保障可复现性,本文还发布了人工标注的基准数据、所有训练模型及完整数据生成流程。
原文摘要 · Abstract (English)
Vision-language models hold considerable promise for ophthalmology, but their development depends on large-scale, high-quality image-text datasets that remain scarce. We present PubMed-Ophtha, a hierarchical dataset of 102,023 ophthalmological image-caption pairs extracted from 15,842 open-access articles in PubMed Central. Unlike existing datasets, figures are extracted directly from article PDFs at full resolution and decomposed into their constituent panels, panel identifiers, and individual images. Each image is annotated with its imaging modality -- color fundus photography, optical coherence tomography, retinal imaging, or other -- and a mark status indicating the presence of annotation marks such as arrows. Figure captions are split into panel-level subcaptions using a two-step LLM approach, achieving a mean average sentence BLEU score of 0.913 on human-annotated data. Panel and image detection models reach a [email protected] of 0.909 and 0.892, respectively, and figure extraction achieves a median IoU of 0.997. To support reproducibility, we additionally release the human-annotated ground-truth data, all trained models, and the full dataset generation pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。