arXiv:2602.05374cs.CLcs.LG2026-02中稿 · HeaLing-EACL 2026被引 2

对比阿拉伯语与英语医学问答模型表现,发现语言差异导致性能差距扩大。

Cross-Lingual Empirical Evaluation of Large Language Models for Arabic Medical Tasks

  • 跨语言对比分析阿拉伯语与英语医学问答任务表现
  • 复杂任务下阿拉伯语模型性能显著下降,差距加剧
  • 模型自信度与正确性关联弱,适合医疗领域多语言研究者

近年来,大语言模型(LLMs)在临床决策支持、医学教育和医学问答等医疗应用中得到广泛应用。然而,这些模型大多以英语为中心,限制了其在语言多样性群体中的鲁棒性和可靠性。已有研究指出,在多种医疗任务中,低资源语言存在性能差异,但根本原因尚不明确。本研究对阿拉伯语和英语的医学问答任务进行了跨语言实证分析。结果表明,随着任务复杂度提升,语言驱动的性能差距持续存在并加剧。分词分析显示阿拉伯语医疗文本存在结构碎片化问题,可靠性分析则表明模型报告的置信度与解释与正确性相关性较弱。这些发现凸显了在医疗任务中设计和评估语言感知型大模型的重要性。

原文摘要 · Abstract (English)

In recent years, Large Language Models (LLMs) have become widely used in medical applications, such as clinical decision support, medical education, and medical question answering. Yet, these models are often English-centric, limiting their robustness and reliability for linguistically diverse communities. Recent work has highlighted discrepancies in performance in low-resource languages for various medical tasks, but the underlying causes remain poorly understood. In this study, we conduct a cross-lingual empirical analysis of LLM performance on Arabic and English medical question and answering. Our findings reveal a persistent language-driven performance gap that intensifies with increasing task complexity. Tokenization analysis exposes structural fragmentation in Arabic medical text, while reliability analysis suggests that model-reported confidence and explanations exhibit limited correlation with correctness. Together, these findings underscore the need for language-aware design and evaluation strategies in LLMs for medical tasks.

大语言模型多语言医疗AI阿拉伯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。