arXiv:2606.03693cs.CLcs.CV2026-06中稿 · CVPR被引 1

测试印尼语对医学视觉问答模型的冲击,发现英语表现好不等于多语言也强。

Does Language Shift Break Medical Vision-Language Models? Indonesian Radiology Visual Question Answering Case Study

论文配图:Does Language Shift Break Medical Vision-Language Models? Indonesian Radiology Visual Question Answering Case Study
图 1 · 摘自论文原文
  • 构建印尼语放射科视觉问答数据集IndoRad-VQA,确保术语和答案一致。
  • 模型在印尼语下准确率下降8%至25%,语言鲁棒性差距明显。
  • 揭示了模型在是非判断、位置识别上的常见错误,适合多语言医疗研究者参考。

医学视觉语言模型(VLMs)通常在英文放射科视觉问答基准上评估,对其在非英文临床语言下的鲁棒性研究不足。本文构建IndoRad-VQA,即VQA-RAD的印尼语版本,评估当问题用巴厘语提出时,医学VLM是否仍具备放射科推理能力。通过基于自评估的质量控制,将放射科问答对翻译为印尼语,保持临床意义、术语一致性及答案等价性。在英文与印尼语提示设置下,评估通用型、东南亚多语言及医学专用型VLMs。除准确率外,还量化了英印尼输入间的语言鲁棒性差距。错误分析揭示了常见失败模式,如是/否颠倒、解剖侧别错误及输出语言不匹配。结果表明,英语医学VQA表现优异并不意味着在印尼语临床场景中同样稳健。不同评估指标下,性能差距达8%至25%。研究呼吁更包容的多语言医学多模态基础模型评估体系。数据集已公开于https://huggingface.co/datasets/Lab-IS/IndoRad-VQA。

原文摘要 · Abstract (English)

Medical Vision-Language Models (VLMs) are typically evaluated on English radiology visual question answering benchmarks, leaving their robustness under non-English clinical language largely unexplored. We introduce IndoRad-VQA, an Indonesian adaptation of VQA-RAD, to assess whether medical VLMs retain radiology reasoning ability when questions are asked in Bahasa Indonesia. Radiology question-answer pairs are translated into Indonesian with self-evaluation-based quality control to preserve clinical meaning, terminology consistency, and answer equivalence. We evaluate general-purpose, Southeast Asian multilingual, and medical-specific VLMs under English and Indonesian prompting settings. Beyond accuracy, we quantify the language robustness gap between English and Indonesian inputs. We also conduct an error analysis to identify failure modes of question answering, such as yes/no flips, laterality errors, and output-language mismatches. Our findings show that strong performance on English medical VQA benchmarks does not necessarily translate to robust behavior in Indonesian clinical contexts. We observe a performance gap of 8 to 25 percent between the English and Indonesian settings, depending on the evaluation metric. These results highlight the need for more inclusive multilingual evaluation of medical multimodal foundation models. The dataset is available at https://huggingface.co/datasets/Lab-IS/IndoRad-VQA.

多语言医学AI视觉问答印尼语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。