首个波兰医学视觉问答基准,揭示模型过度依赖文本而非图像。
Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence

- 构建波兰医学VQA数据集,含多领域图像与文本题
- 最佳模型准确率79.0%,多数仍低于人类表现
- 模型更依赖问题文本,图像重要性被低估
我们引入首个波兰语医学视觉问答(VQA)基准,基于波兰执业医师和牙医专科认证考试题目构建。该基准包含跨多个医学专业和视觉领域的带图问题,以及仅含文本的问题回答对照集。评估了面向波兰语、通用开源及商用视觉语言模型。任务仍具挑战:最佳模型在完整VQA集上达79.0%准确率,仅GPT-5.6在有候选答案的子集上超过近似人类参考;其余模型均低于人类表现。为评估视觉定位能力,对比完整输入与省略图像、问题或两者的情况,并按图像重要性分类。结果显示模型从问题文本中获取的信息多于图像,在图像主导型问题上表现更差。在问答与视觉问答任务中,模型仅凭答案选项即可获得显著高于随机水平的准确率,表明关键任务组件缺失时仍可维持非平凡性能。
原文摘要 · Abstract (English)
We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists pursuing specialist certification. The benchmark comprises image-containing questions spanning diverse medical specialties and visual domains, together with a text-only question answering (QA) control set. We evaluate Polish-oriented, general-purpose open-weight, and commercial vision-language models. The task remains challenging: the best model achieves 79.0\% accuracy on the full VQA set, and only GPT-5.6 surpasses the approximate human reference on the subset with available candidate responses; all other evaluated models perform worse than humans. To assess visual grounding, we compare complete inputs with configurations omitting the image, the question, or both, and categorize questions by image importance. Models derive more useful information from the question text than from the image and perform worse on image-dominant questions. Across both QA and VQA, they nevertheless achieve above-chance accuracy from the answer choices alone, showing that non-trivial performance can persist even when key task components are missing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。