首个涵盖10类日本医疗执照考试的多模态评测基准,检验大模型临床推理能力。
KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations
- 构建包含11588道真实考题的多模态数据集,融合临床图像与专家解析。
- 30多个主流大模型在文本和图像任务中均未全领域达标,最高通过率不足60%。
- 专为多语言、多专业医疗场景设计,适合医疗AI系统评估与改进。
大语言模型在医学执照考试中已展现显著性能,但对多种医疗角色、高风险临床场景的综合评估仍具挑战。现有基准多为纯文本、以英文为主,且聚焦药物知识,难以评估广泛医疗知识与多模态推理能力。为此,我们推出KokushiMD-10,首个基于日本10项国家级医疗执照考试构建的多模态基准。该数据集覆盖医学、牙科、护理、药学及辅助医疗等多个领域,包含超过11588道真实考题,整合临床图像与专家标注的解题逻辑,用于评估文本与视觉推理能力。我们在文本与图像双设置下对30余款先进大模型(如GPT-4o、Claude 3.5、Gemini)进行评测。尽管表现亮眼,但无一模型能在所有领域稳定通过考试,凸显医疗AI在推理能力上的持续挑战。KokushiMD-10为多语言、多模态临床任务中的推理导向型医疗AI提供了全面且语言精准的评估资源。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have demonstrated notable performance in medical licensing exams. However, comprehensive evaluation of LLMs across various healthcare roles, particularly in high-stakes clinical scenarios, remains a challenge. Existing benchmarks are typically text-based, English-centric, and focus primarily on medicines, which limits their ability to assess broader healthcare knowledge and multimodal reasoning. To address these gaps, we introduce KokushiMD-10, the first multimodal benchmark constructed from ten Japanese national healthcare licensing exams. This benchmark spans multiple fields, including Medicine, Dentistry, Nursing, Pharmacy, and allied health professions. It contains over 11588 real exam questions, incorporating clinical images and expert-annotated rationales to evaluate both textual and visual reasoning. We benchmark over 30 state-of-the-art LLMs, including GPT-4o, Claude 3.5, and Gemini, across both text and image-based settings. Despite promising results, no model consistently meets passing thresholds across domains, highlighting the ongoing challenges in medical AI. KokushiMD-10 provides a comprehensive and linguistically grounded resource for evaluating and advancing reasoning-centric medical AI across multilingual and multimodal clinical tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。