构建多语言多模态医疗问答数据集,评估AI在真实医疗场景中的表现
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation
- 收录4国568个带医学图像的多选题,覆盖原生语言与临床验证翻译
- 测试模型在有无图像、不同语言下的表现,发现图文结合提升准确率
- 适合医疗AI开发者和评估者,推动跨文化公平性研究
多模态视觉语言模型(VLMs)正日益应用于全球医疗场景,亟需可靠基准来确保其安全性、有效性与公平性。现有医学多选题数据集多为纯文本,且语言和国家覆盖有限。为此,我们推出WorldMedQA-V——一个更新的多语言、多模态基准数据集,用于评估医疗领域VLMs。该数据集包含来自巴西、以色列、日本和西班牙的568个标注多选题及其对应的568张医学图像,涵盖原始语言及由母语临床医生验证的英文翻译。提供主流开源与闭源模型在本地语言与英文翻译、有图与无图条件下的基线性能。该基准旨在更贴近AI实际部署的多样化医疗环境,促进更具公平性、有效性和代表性的应用。
原文摘要 · Abstract (English)
Multimodal/vision language models (VLMs) are increasingly being deployed in healthcare settings worldwide, necessitating robust benchmarks to ensure their safety, efficacy, and fairness. Multiple-choice question and answer (QA) datasets derived from national medical examinations have long served as valuable evaluation tools, but existing datasets are largely text-only and available in a limited subset of languages and countries. To address these challenges, we present WorldMedQA-V, an updated multilingual, multimodal benchmarking dataset designed to evaluate VLMs in healthcare. WorldMedQA-V includes 568 labeled multiple-choice QAs paired with 568 medical images from four countries (Brazil, Israel, Japan, and Spain), covering original languages and validated English translations by native clinicians, respectively. Baseline performance for common open- and closed-source models are provided in the local language and English translations, and with and without images provided to the model. The WorldMedQA-V benchmark aims to better match AI systems to the diverse healthcare environments in which they are deployed, fostering more equitable, effective, and representative applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。