arXiv:2604.14306cs.CLcs.AI2026-04

构建首个多语言多模态医学考试数据集,评估大模型跨语言医疗推理能力。

EuropeMedQA Study Protocol: A Multilingual, Multimodal Medical Examination Dataset for Language Model Evaluation

  • 基于欧陆四国官方医学考试,构建多语言多模态数据集
  • 采用零样本约束提示策略,验证模型跨语言迁移与视觉推理能力
  • 面向医疗AI研发者,助力提升模型在真实临床场景的泛化性

尽管大语言模型在英语医学考试中表现优异,但在非英语语言和多模态诊断任务上性能常显著下降。本研究协议描述了EuropeMedQA的构建,这是首个源自意大利、法国、西班牙和葡萄牙官方监管考试的综合性、多语言、多模态医学考试数据集。遵循FAIR数据原则和SPIRIT-AI指南,我们制定了严格的筛选流程与自动化翻译管道,用于对比分析。采用零样本、严格约束的提示策略,评估当代多模态大模型在跨语言迁移与视觉推理方面的能力。EuropeMedQA旨在提供一个抗污染基准,反映欧洲临床实践的复杂性,推动更通用医疗人工智能的发展。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) have demonstrated high proficiency on English-centric medical examinations, their performance often declines when faced with non-English languages and multimodal diagnostic tasks. This study protocol describes the development of EuropeMedQA, the first comprehensive, multilingual, and multimodal medical examination dataset sourced from official regulatory exams in Italy, France, Spain, and Portugal. Following FAIR data principles and SPIRIT-AI guidelines, we describe a rigorous curation process and an automated translation pipeline for comparative analysis. We evaluate contemporary multimodal LLMs using a zero-shot, strictly constrained prompting strategy to assess cross-lingual transfer and visual reasoning. EuropeMedQA aims to provide a contamination-resistant benchmark that reflects the complexity of European clinical practices and fosters the development of more generalizable medical AI.

医学AI多语言多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。