首个韩语医学影像多模态问答基准,评测模型看图识病能力。
KorMedMCQA-V: A Multimodal Benchmark for Evaluating Vision-Language Models on the Korean Medical Licensing Examination
- 构建1534道韩医执照考题多模态数据集,含2043张医学影像。
- 顶尖模型准确率达96.9%,但跨图像推理题性能普遍下降。
- 适合医疗AI、多模态模型研究者评估韩国医学视觉理解能力。
我们提出KorMedMCQA-V,一个面向韩语医学执照考试的多模态多选题问答基准,用于评估视觉-语言模型(VLMs)在医学影像理解方面的能力。该数据集包含2012至2023年韩医执照考试中的1,534道题目,关联2,043张医学影像,约30%题目需整合多幅图像进行推理。影像涵盖X射线、计算机断层扫描(CT)、心电图(ECG)、超声、内窥镜及其他临床视觉内容。我们在统一零样本评估协议下测试了超过50个模型,包括通用型、医学专用型及韩语专用型。最佳商用模型(Gemini-3.0-Pro)准确率达96.9%,最优开源模型(Qwen3-VL-32B-Thinking)为83.7%,而最佳韩语专用模型(VARCO-VISION-2.0-14B)仅为43.2%。发现推理导向模型比指令微调版本最高提升20个百分点,医学领域专业化对强通用基线表现不一致,所有模型在多图像问题上性能均下降,且不同成像模态间表现差异显著。KorMedMCQA-V与文本仅有的KorMedMCQA形成完整评估体系,覆盖文本与多模态场景。数据集已通过Hugging Face Datasets发布:https://huggingface.co/datasets/seongsubae/KorMedMCQA-V。
原文摘要 · Abstract (English)
We introduce KorMedMCQA-V, a Korean medical licensing-exam-style multimodal multiple-choice question answering benchmark for evaluating vision-language models (VLMs). The dataset consists of 1,534 questions with 2,043 associated images from Korean Medical Licensing Examinations (2012-2023), with about 30% containing multiple images requiring cross-image evidence integration. Images cover clinical modalities including X-ray, computed tomography (CT), electrocardiography (ECG), ultrasound, endoscopy, and other medical visuals. We benchmark over 50 VLMs across proprietary and open-source categories-spanning general-purpose, medical-specialized, and Korean-specialized families-under a unified zero-shot evaluation protocol. The best proprietary model (Gemini-3.0-Pro) achieves 96.9% accuracy, the best open-source model (Qwen3-VL-32B-Thinking) 83.7%, and the best Korean-specialized model (VARCO-VISION-2.0-14B) only 43.2%. We further find that reasoning-oriented model variants gain up to +20 percentage points over instruction-tuned counterparts, medical domain specialization yields inconsistent gains over strong general-purpose baselines, all models degrade on multi-image questions, and performance varies notably across imaging modalities. By complementing the text-only KorMedMCQA benchmark, KorMedMCQA-V forms a unified evaluation suite for Korean medical reasoning across text-only and multimodal conditions. The dataset is available via Hugging Face Datasets: https://huggingface.co/datasets/seongsubae/KorMedMCQA-V.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。