arXiv:2604.20531cs.CL2026-04

跨语言证据能提升低资源医学问答,但效果依赖语言和模型大小。

Effects of Cross-lingual Evidence in Multilingual Medical Question Answering

论文配图:Effects of Cross-lingual Evidence in Multilingual Medical Question Answering
图 1 · 摘自论文原文
  • 用多语言、单语言和跨语言检索结合外部知识源。
  • 英语高资源语言用网页数据提升性能,低资源语言需双语检索。
  • 大模型在英语上表现更好,但跨语言策略对低资源语言更关键。

本文研究了在高资源语言(英语、西班牙语、法语、意大利语)和低资源语言(巴斯克语、哈萨克语)之间的多语言医学问答。评估了三种外部证据来源:专业医学知识库、网络检索内容以及大模型的参数化知识。实验涵盖多语言、单语言和跨语言检索。结果表明,大模型在英语基准测试中始终表现更优。引入外部知识后,英语网页数据对高资源语言最有效;而对低资源语言,结合英语与目标语言检索的策略可达到与高资源语言相当的准确率。该发现挑战了外部知识普遍提升性能的假设,揭示有效策略取决于语言资源来源和模型规模。此外,如PubMed等专业医学知识库虽权威,但缺乏充分的多语言覆盖。

原文摘要 · Abstract (English)

This paper investigates Multilingual Medical Question Answering across high-resource (English, Spanish, French, Italian) and low-resource (Basque, Kazakh) languages. We evaluate three types of external evidence sources across models of varying size: curated repositories of specialized medical knowledge, web-retrieved content, and explanations from LLM's parametric knowledge. Moreover, we conduct experiments with multilingual, monolingual and cross-lingual retrieval. Our results demonstrate that larger models consistently achieve superior performance in English across baseline evaluations. When incorporating external knowledge, web-retrieved data in English proves most beneficial for high-resource languages. Conversely, for low-resource languages, the most effective strategy combines retrieval in both English and the target language, achieving comparable accuracy to high-resource language results. These findings challenge the assumption that external knowledge systematically improves performance and reveal that effective strategies depend on both the source of language resources and on model scale. Furthermore, specialized medical knowledge sources such as PubMed are limited: while they provide authoritative expert knowledge, they lack adequate multilingual coverage

医学问答多语言跨语言低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。