arXiv:2508.16390cs.CLcs.AI2025-08中稿 · npj Digital Medici…被引 5

首个罗马尼亚语医学问答基准,评估大模型在癌症诊疗中的表现。

A Large-Scale Benchmark for Evaluating Large Language Models on Medical Question Answering in Romanian

  • 构建10.6万条罗马尼亚语医学问答数据,涵盖1242名患者病历。
  • 微调模型性能远超零样本提示,说明预训练模型无法直接泛化。
  • 适合关注多语言医疗AI、临床决策支持的开发者与研究者。

我们提出MedQARo,首个针对罗马尼亚语的大型医学问答基准,并对前沿大语言模型(LLMs)进行综合评估。数据集由两家医疗中心提供的1,242名癌症患者病例生成,包含105,880个高质量问答对,问题涉及病例摘要,需关键词提取与推理能力。基准包含域内及跨域测试集(跨中心、跨癌种),可精准评估模型泛化能力。我们在MedQARo上测试了四种来自不同架构的开源模型,每种模型均在零样本提示和监督微调两种场景下运行;同时评估了通过API访问的两个先进模型:GPT-5.2和Gemini 3 Flash。结果表明,微调模型显著优于零样本模型,说明预训练模型在该任务上难以泛化。研究强调了领域特定与语言特定微调对罗马尼亚语临床问答可靠性的重要性。

原文摘要 · Abstract (English)

We introduce MedQARo, the first large-scale medical QA benchmark in Romanian, alongside a comprehensive evaluation of state-of-the-art large language models (LLMs). We construct a high-quality and large-scale dataset comprising 105,880 QA pairs about cancer patients from two medical centers. The questions regard medical case summaries of 1,242 patients, requiring both keyword extraction and reasoning. Our benchmark contains both in-domain and cross-domain (cross-center and cross-cancer) test collections, enabling a precise assessment of generalization capabilities. We experiment with four open-source LLMs from distinct families of models on MedQARo. Each model is employed in two scenarios: zero-shot prompting and supervised fine-tuning. We also evaluate two state-of-the-art LLMs exposed only through APIs, namely GPT-5.2 and Gemini 3 Flash. Our results show that fine-tuned models significantly outperform zero-shot models, indicating that pretrained models fail to generalize on MedQARo. Our findings demonstrate the importance of both domain-specific and language-specific fine-tuning for reliable clinical QA in Romanian.

医学问答多语言大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。