解决多语言医疗视觉问答中的性能下降问题,提升非英语表现。
Analyzing and Mitigating Cross-Lingual Degradation in Multilingual Medical VQA

- 构建八语言医疗VQA基准,分四类场景隔离核心能力
- 发现跨语言退化程度因任务场景而异,非均匀分布
- 提出无需训练的方法,在推理时对齐非英语表示
医疗视觉问答(VQA)是临床AI中的关键任务,但现有评估几乎仅针对英语,难以覆盖多语言患者和医生。近期多语言医疗VQA基准显示,大型视觉语言模型(LVLMs)在非英语语言中性能下降,但缺乏对跨语言差异如何影响医疗VQA关键能力的细粒度分析。为此,我们构建了一个覆盖八种语言的多语言医疗VQA基准,按四个代表性场景组织,以分离出医疗VQA所需的核心能力。评估五种开源与闭源LVLMs发现,跨语言退化并非均匀发生,而是高度依赖于具体场景。因此,我们提出MedVL-XLRepE——一种无需训练的、场景感知的表示工程方法,利用模型在英语医疗VQA上的优异表现,在推理阶段引导非英语表示向其英语对应项对齐。在三种LVLM和八种语言上,MedVL-XLRepE持续缓解跨语言退化,最高提升6.33%。
原文摘要 · Abstract (English)
Medical visual question answering (VQA) is a crucial task in clinical AI, yet its evaluation has so far centered almost exclusively on English, limiting its relevance to linguistically diverse patients and clinicians. Recent multilingual medical VQA benchmarks show that large vision-language models (LVLMs) degrade in non-English languages, but lack a fine-grained analysis of how cross-lingual variation affects the distinct capabilities that medical VQA requires. To this end, we construct a multilingual medical VQA benchmark over eight languages, organized into four representative scenarios that isolate the core capabilities medical VQA requires. Evaluating five open- and closed-source LVLMs, we find that cross-lingual degradation is not uniform but highly scenario-dependent. We therefore propose MedVL-XLRepE, a training-free scenario-aware representation engineering method, leveraging LVLMs' superior English medical VQA capability to steer non-English representations toward their English counterparts at inference time. Across three LVLMs and eight languages, MedVL-XLRepE consistently mitigates cross-lingual degradation, with gains of up to 6.33\%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。