首个孟加拉语医学视觉问答数据集,揭示大模型在本地化医疗场景表现欠佳
How Good LLMs Are at Answering Bangla Medical Visual Questions? Dataset and Benchmarking

- 构建孟加拉语临床图像问答数据集BanglaMedVQA
- GPT-4.1 mini等顶级模型在诊断题上准确率仍很低
- 适合关注低资源语言医疗AI的研究者参考
大型语言模型(LLMs)和大型视觉语言模型(LVLMs)的进展使通用系统在复杂推理任务中展现出潜力,尤其在医学领域。医学视觉问答(MedVQA)受益显著。然而,尽管孟加拉语是全球使用最广泛的语言之一,目前尚无相关基准。为此,我们推出了BanglaMedVQA,一个包含临床验证的图像-问题-答案对的数据集,并对当前基础模型进行了全面评估。与英文MedVQA基准上模型表现不佳的发现一致,本研究显示孟加拉语性能显著偏低,反映出低资源语言固有的挑战。即使是表现最佳的Gemini和GPT-4.1 mini也难以准确回答专业诊断问题,表明其在细粒度医学推理方面存在严重局限。尽管某些开源模型如Gemma-3在一般类别中偶尔表现更优,但在临床复杂问题上同样表现困难,凸显了高质量评估方法的紧迫需求。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) and Large Vision Language Models (LVLMs) have enabled general-purpose systems to demonstrate promising capabilities in complex reasoning tasks, including those in the medical domain. Medical Visual Question Answering (MedVQA) has particularly benefited from these developments. However, despite Bangla being one of the most widely spoken languages globally, there exists no established MedVQA benchmark for it. To address this gap, we introduce BanglaMedVQA, a dataset comprising clinically validated image-question-answer pairs, along with a comprehensive evaluation of current foundation models on this resource. Consistent with prior findings that report low performance of current models on English MedVQA benchmarks, our analysis reveals that Bangla performance is substantially lower, reflecting the challenges inherent to low-resource languages. Even top-performing models such as Gemini and GPT-4.1 mini fail to accurately answer specialized diagnostic questions, indicating severe limitations in fine-grained medical reasoning. Although certain open-source models, such as Gemma-3, occasionally outperform these models in general categories, they too struggle with clinically complex questions, underscoring the urgent need for top-notch evaluation method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。