多语言医学大模型评估基准,揭示模型在低资源语言中表现差异
HealMed: Multilingual Evaluation of Large Language Models in Medicine

- 构建覆盖9种语言的医学多任务评测集,含1000例专家审核样本
- 低资源语言下模型性能下降明显,开源与专用模型差距不一
- 翻译质量直接影响评估结果,强调跨语言评测需人工校验
我们提出 HealMed,一个面向医学领域的大语言模型多语言评估基准。该基准包含九种语言各1000个样本,源自九个数据集,涵盖三种任务类型:多项选择题问答(MCQA)、自然语言推理(NLI)和开放问答。基准历时两年,由来自九个国家和地区的23名医生和医学专家共同开发。每项翻译均由两名精通英语及目标语言的专家评审修订。在 HealMed 上,模型性能在低资源语言中下降最显著,但不同语言和模型间的差距差异明显。最强的专有模型在跨语言间表现最稳定,而许多开源和医学专用模型则表现出更大且不一致的性能差距。医学专业化本身并不能保证多语言鲁棒性。此外,专家修订可能提升或降低测评性能,表明翻译质量对跨语言评估结果有重大影响。
原文摘要 · Abstract (English)
We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn from nine datasets and covering three task formats: MCQA, NLI and open-ended QA. The benchmark was developed over two years by 23 physicians and medical experts based across nine countries and regions. Each translation was evaluated and revised by two experts fluent in English and the corresponding target language. On HealMed, performance declined most in low-resource languages, although the size of the gap varied markedly across languages and models. The strongest proprietary models were the most stable across languages, whereas many open-source and medically specialized models showed larger and less consistent gaps. Medical specialization alone did not ensure multilingual robustness. Furthermore, expert revision could either raise or lower measured performance, indicating that translation quality materially affects cross-language evaluation results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。