首个葡萄牙语医疗AI评估基准,揭示大模型在护理等专业上的知识短板。
HealthQA-BR: A System-Wide Benchmark Reveals Critical Knowledge Gaps in Large Language Models

- 构建涵盖多专业的真实医疗考试题库,评估大模型系统性知识
- 顶尖模型整体准确率86.6%,但社会工作仅68.4%、神经外科仅60.0%
- 首次暴露大模型知识分布不均问题,适合关注医疗AI安全的从业者
当前医疗大模型评估多依赖以医生为中心的英文基准,造成能力误判。我们推出HealthQA-BR——首个针对葡萄牙语医疗场景的大规模系统级评测基准,包含巴西国家执照与住院医师考试中的5,632道题目,覆盖医学、护理、牙科、心理学、社会工作等多专业。对超过20个主流大模型进行零样本评估发现,尽管最先进的GPT-4.1总体准确率达86.6%,但知识表现呈明显“尖峰”特征:眼科达98.7%,神经外科仅60.0%,社会工作为68.4%。这一跨模型普遍现象表明,高平均分无法反映真实风险。通过公开数据集与评估工具,推动从单一指标转向全面、细致的医疗AI能力审计。
原文摘要 · Abstract (English)
The evaluation of Large Language Models (LLMs) in healthcare has been dominated by physician-centric, English-language benchmarks, creating a dangerous illusion of competence that ignores the interprofessional nature of patient care. To provide a more holistic and realistic assessment, we introduce HealthQA-BR, the first large-scale, system-wide benchmark for Portuguese-speaking healthcare. Comprising 5,632 questions from Brazil's national licensing and residency exams, it uniquely assesses knowledge not only in medicine and its specialties but also in nursing, dentistry, psychology, social work, and other allied health professions. We conducted a rigorous zero-shot evaluation of over 20 leading LLMs. Our results reveal that while state-of-the-art models like GPT 4.1 achieve high overall accuracy (86.6%), this top-line score masks alarming, previously unmeasured deficiencies. A granular analysis shows performance plummets from near-perfect in specialties like Ophthalmology (98.7%) to barely passing in Neurosurgery (60.0%) and, most notably, Social Work (68.4%). This "spiky" knowledge profile is a systemic issue observed across all models, demonstrating that high-level scores are insufficient for safety validation. By publicly releasing HealthQA-BR and our evaluation suite, we provide a crucial tool to move beyond single-score evaluations and toward a more honest, granular audit of AI readiness for the entire healthcare team.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。