arXiv:2606.25246cs.CVcs.CL2026-06

构建首个中英双语血细胞视觉问答数据集,助力多语言医疗AI发展

Multilingual Hematology Visual Question Answering Dataset

论文配图:Multilingual Hematology Visual Question Answering Dataset
图 1 · 摘自论文原文
  • 基于形态学标注构建双语医学图像问答数据集
  • 包含11万对中英文问答,覆盖2万张白细胞图像
  • 专为巴基斯坦等多语言地区医疗场景设计

视觉语言模型(VLMs)在医学图像分析中展现出巨大潜力,但现有血液病视觉语言资源仍以英语为主,难以满足多语言医疗环境需求。针对南亚及巴基斯坦等地普遍使用乌尔都语却依赖英文医疗系统的问题,我们通过调研发现临床文档与患者沟通间存在显著语言错配。为此,提出WBCMor VQA——一个经临床验证的中英双语、形态感知的白细胞分析视觉问答基准。该数据集基于LeukemiaAttri和WBCAtt的形态学标注,并结合领域专用乌尔都语血液病词典,确保语言一致性和临床准确性。最终包含11万对双语问答,对应2万张白血病与正常单细胞图像。我们评估了多个开源VLM在该基准上的表现,旨在推动可访问、临床相关的多语言医疗AI系统发展。

原文摘要 · Abstract (English)

Vision Language Models (VLMs) have shown promising capabilities in medical image analysis by jointly understanding visual and textual information for tasks such as Visual Question Answering. However, existing hematology vision-language resources remain predominantly English centric, limiting their applicability in multilingual healthcare environments. This challenge is releveant generally to South Asia and specifically to Pakistan, where Urdu is widely used despite healthcare information and digital medical systems being largely dependent on English. To investigate this gap, we conducted a survey among healthcare professionals, which revealed substantial language mismatches between clinical documentation and patient communication, emphasizing the need for multilingual healthcare technologies. To address this limitation, we introduce WBCMor VQA, a clinically validated bilingual English, Urdu morphology aware VQA benchmark for leukemia and normal white blood cell analysis. The benchmark is constructed using morphology-aware annotations from LeukemiaAttri and WBCAtt datasets and supported by a domain specific Urdu hematology dictionary to ensure linguistic consistency and clinical correctness. The final benchmark contains 110K bilingual question answer pairs serving as VQA annotations for 20K leukemic and normal single-cell images. Furthermore, we establish baseline performance by evaluating multiple open-source VLMs on the proposed benchmark. The proposed resource aims to facilitate the development of accessible and clinically relevant AI systems for multilingual healthcare environments.

视觉问答多语言医疗AI血液病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。