首个波兰语医学考试基准,测试大模型跨语言医疗能力
Polish-English medical knowledge transfer: A new benchmark and results
- 构建2.4万+题波兰医学考试数据集,含中英双语对照
- GPT-4o接近人类水平,但跨语言理解仍存短板
- 揭示多语言医疗AI差距,适合临床决策研究者参考
大型语言模型在专业任务中展现巨大潜力,尤其在医疗问题求解方面。然而,现有研究主要聚焦英语场景。本研究基于波兰医学执照与专科考试(LEK、LDEK、PES)构建全新基准数据集,涵盖超过24,000道试题,其中包含由考试中心为外籍考生专业翻译的波兰语-英语平行语料。数据来自医学考试中心与首席医疗委员会公开资源。通过该结构化基准,系统评估通用、领域专用及波兰语专用模型表现,并与医学生水平对比。结果显示,尽管GPT-4o已接近人类表现,但在跨语言翻译与专科理解方面仍存在显著挑战。研究揭示了模型在不同语言和医学专科间的性能差异,强调将大模型应用于临床实践中的局限性与伦理考量。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated significant potential in handling specialized tasks, including medical problem-solving. However, most studies predominantly focus on English-language contexts. This study introduces a novel benchmark dataset based on Polish medical licensing and specialization exams (LEK, LDEK, PES) taken by medical doctor candidates and practicing doctors pursuing specialization. The dataset was web-scraped from publicly available resources provided by the Medical Examination Center and the Chief Medical Chamber. It comprises over 24,000 exam questions, including a subset of parallel Polish-English corpora, where the English portion was professionally translated by the examination center for foreign candidates. By creating a structured benchmark from these existing exam questions, we systematically evaluate state-of-the-art LLMs, including general-purpose, domain-specific, and Polish-specific models, and compare their performance against human medical students. Our analysis reveals that while models like GPT-4o achieve near-human performance, significant challenges persist in cross-lingual translation and domain-specific understanding. These findings underscore disparities in model performance across languages and medical specialties, highlighting the limitations and ethical considerations of deploying LLMs in clinical practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。