arXiv:2606.07853cs.CLcs.AI2026-06被引 1

首个葡萄牙语临床大模型评估基准,揭示语言差距因任务而异。

Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese

论文配图:Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese
图 1 · 摘自论文原文
  • 构建首个双语临床案例库,含2892例真实病例。
  • 诊断检索英语优势明显,其他任务中葡语表现相当甚至略优。
  • 适合关注多语言医疗AI、跨语言公平性研究者阅读。

大型语言模型正重塑临床决策支持,但现有评测多基于英语,存在语言鸿沟。本文推出ClinicalBr,首个基于真实巴西病例报告的双语临床评测基准,涵盖2892例来自28本SciELO医学期刊的病例,覆盖18个专科,以平行葡萄牙语-英语对形式呈现。每个病例支持四项任务:诊断检索、鉴别诊断、检查建议与治疗规划。评估MedGemma-27B、Sabiá-4、DeepSeek-R1和o3-mini四款模型在双语下的表现。结果显示,语言差距依赖任务:诊断检索中英语普遍领先7.5–12.1个百分点;而在鉴别诊断、检查建议与治疗规划中,多数模型置信区间包含零,葡萄牙语完成度甚至略高。巴西特有疾病表现优于全集,表明热带病种已充分覆盖于预训练数据。检查建议为最难任务,所有模型F1均低于0.10,远低于鉴别诊断上限(0.20–0.27)。

原文摘要 · Abstract (English)

Large Language Models are transforming the support for clinical decision and their application in real scenarios. Yet, most benchmarks are conducted in English, and cross-lingual evaluation is needed to tackle the language gaps in global access. We introduce ClinicalBr, the first bilingual benchmark for clinical decision built from real Brazilian case reports. The corpus contains 2,892 cases drawn from 28 SciELO medical journals, spanning 18 specialties, and is structured as parallel Portuguese-English pairs. Each case supports four evaluation tasks: diagnosis retrieval, differential diagnosis, exam recommendation, and treatment planning. We evaluate four models: MedGemma-27B, Sabiá-4, DeepSeek-R1, and o3-mini, across both languages. The central finding is that the Portuguese-English performance gap is task-dependent, not general. In diagnosis retrieval, English yields a consistent advantage across all models, with +7.5-12.1 accuracy points. This advantage disappears in differential diagnosis, exam recommendation, and treatment planning, where confidence intervals cross zero for most models and Portuguese completeness scores are marginally higher. Brazilian-endemic conditions proved easier than the full corpus, not harder, indicating that tropical presentations are adequately represented in current pre-training. Exam recommendation was the hardest task across all models and both languages, with F1 scores below 0.10, well below the differential diagnosis ceiling of 0.20-0.27.

临床AI多语言评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。