首个眼科多模态问答基准,评估大模型在真实临床场景下的表现
Benchmarking Large Multimodal Models for Ophthalmic Visual Question Answering with OphthalWeChat
- 基于微信公众号真实眼科图像与双语问答对构建评测数据集
- Gemini 2.0 Flash整体准确率54.8%领先,中文/英文均表现最优
- 涵盖68种影像组合,适配医生、AI研发者及医学影像研究者
目的:构建用于评估眼科领域视觉问答(VQA)大模型的双语基准。方法:从微信公众号收集2016年1月1日至2024年12月31日发布的眼科图像及其图文内容,利用GPT-4o-mini生成中英文问答对,按问题类型与语言划分为六类:二分类(Binary_CN, Binary_EN)、单选题(Single-choice_CN, Single-choice_EN)和开放式(Open-ended_CN, Open-ended_EN)。使用该基准评估GPT-4o、Gemini 2.0 Flash和Qwen2.5-VL-72B-Instruct三款视觉语言模型。结果:最终的OphthalWeChat数据集包含3,469张图像、30,120组问答对,覆盖9个眼科亚专科、548种疾病、29种成像模态及68种模态组合。Gemini 2.0 Flash总体准确率达0.548,显著高于GPT-4o(0.522,P < 0.001)和Qwen2.5-VL-72B-Instruct(0.514,P < 0.001),在中文(0.546)和英文(0.550)子集均领先。其在Binary_CN(0.687)、Single-choice_CN(0.666)和Single-choice_EN(0.646)上表现突出;而GPT-4o在Binary_EN(0.717)、Open-ended_CN(BLEU-1: 0.301; BERTScore: 0.382)和Open-ended_EN(BLEU-1: 0.183; BERTScore: 0.240)上最佳。结论:本研究首次提出面向眼科的双语多模态问答基准,具有真实临床背景与患者多模态检查记录,支持对大模型进行量化评估,助力开发精准、专业且可信的眼科AI系统。
原文摘要 · Abstract (English)
Purpose: To develop a bilingual multimodal visual question answering (VQA) benchmark for evaluating VLMs in ophthalmology. Methods: Ophthalmic image posts and associated captions published between January 1, 2016, and December 31, 2024, were collected from WeChat Official Accounts. Based on these captions, bilingual question-answer (QA) pairs in Chinese and English were generated using GPT-4o-mini. QA pairs were categorized into six subsets by question type and language: binary (Binary_CN, Binary_EN), single-choice (Single-choice_CN, Single-choice_EN), and open-ended (Open-ended_CN, Open-ended_EN). The benchmark was used to evaluate the performance of three VLMs: GPT-4o, Gemini 2.0 Flash, and Qwen2.5-VL-72B-Instruct. Results: The final OphthalWeChat dataset included 3,469 images and 30,120 QA pairs across 9 ophthalmic subspecialties, 548 conditions, 29 imaging modalities, and 68 modality combinations. Gemini 2.0 Flash achieved the highest overall accuracy (0.548), outperforming GPT-4o (0.522, P < 0.001) and Qwen2.5-VL-72B-Instruct (0.514, P < 0.001). It also led in both Chinese (0.546) and English subsets (0.550). Subset-specific performance showed Gemini 2.0 Flash excelled in Binary_CN (0.687), Single-choice_CN (0.666), and Single-choice_EN (0.646), while GPT-4o ranked highest in Binary_EN (0.717), Open-ended_CN (BLEU-1: 0.301; BERTScore: 0.382), and Open-ended_EN (BLEU-1: 0.183; BERTScore: 0.240). Conclusions: This study presents the first bilingual VQA benchmark for ophthalmology, distinguished by its real-world context and inclusion of multiple examinations per patient. The dataset reflects authentic clinical decision-making scenarios and enables quantitative evaluation of VLMs, supporting the development of accurate, specialized, and trustworthy AI systems for eye care.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。