arXiv:2605.22080cs.CVcs.AI2026-05

构建日本医疗执照多专业视觉语言评估基准,验证模型对医学图像的利用能力。

JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation

论文配图:JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation
图 1 · 摘自论文原文
  • 基于日本官方考试资料构建4.8万道题库,含2万张标注图像。
  • 发现专有模型依赖图像提升显著,但医学专用模型几乎不使用图像信息。
  • 适合关注医疗AI评测、视觉语言模型在医学场景泛化能力的研究者。

我们提出JMed48k,一个用于评估视觉语言模型的多专业日本医疗执照基准。该数据集源自日本厚生劳动省发布的官方PDF材料,涵盖2005至2025年间11项全国性执照考试,包含48,862道考题和20,142张图像,视觉内容按8类体系标注。从中衍生出JMed48k-Eval,为近五年12,484道评分题的子集,包括9,905道纯文本题与2,579道带图像题。我们评估了21种专有、开源及医学专用模型,分别报告纯文本与含图像性能。由于子集题目不同,我们引入配对图像移除审计,考察有图与去图后的答案变化,揭示四种回答转移状态。结果表明:专有与开源模型在图像加持下表现大幅提升,而医学专用系统对视觉证据使用有限,许多正确答案在移除图像后仍保持。即使在专有模型中,图像移除带来的净影响在不同职业间差异达七倍,从医师题+5.7分至公共卫生护士题+39.8分。我们公开发布JMed48k,以支持可复现、职业分层的医学执照场景下视觉语言模型评估。

原文摘要 · Abstract (English)

We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF materials released by the Japanese Ministry of Health, Labour and Welfare, JMed48k contains 48,862 exam questions and 20,142 images from 11 national licensing examinations between 2005 and 2025, with visual content annotated under an 8-type taxonomy. From this corpus, we derive JMed48k-Eval, a recent five-year evaluation subset with 12,484 scored questions, including 9,905 text-only questions and 2,579 questions with images. We evaluate 21 proprietary, open-source, and medical-specific models, reporting text-only and with-image performance separately. Because these subsets contain different questions, we further introduce a paired image-removal audit that evaluates questions with images before and after removing visual content to explore four answer-transition states. The audit shows that proprietary and open source models gain substantially from images, whereas medical-specific systems show limited observable use of visual evidence, with many correct answers persisting after image removal. Even among proprietary models, the net image-removal effect varies sevenfold across professions, from +5.7 points on Physician questions to +39.8 points on Public Health Nurse questions. We release JMed48k to support reproducible, profession-stratified evaluation of vision-language models in medical licensing settings.

医学AI视觉语言评测基准日本医疗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。