arXiv:2512.17028cs.CLcs.AI2025-12被引 2

首个女性健康大模型评测基准,揭示主流模型60%失败率

A Women's Health Benchmark for Large Language Models

  • 构建涵盖5个专科的女性健康评测基准,覆盖3类问题和8种错误类型
  • 13个顶尖大模型平均60%失败率,尤其在遗漏紧急信号上表现差
  • 适合医疗AI安全评估、政策制定者及关注女性健康的研发人员

随着大型语言模型(LLMs)成为数百万用户获取健康信息的主要来源,其在女性健康领域的准确性仍缺乏系统评估。我们提出首个专门针对女性健康的大模型评测基准——女性健康基准(WHB),包含96个经过严格验证的模型片段,覆盖产科与妇科、急诊医学、初级保健、肿瘤学和神经病学五个医学专科,三种查询类型(患者咨询、医生咨询、证据/政策查询),以及八类错误(用药剂量错误、关键信息缺失、过时指南/治疗建议、错误治疗建议、事实性错误、误诊或漏诊、遗漏紧急情况、不当建议)。我们评估了13个前沿大模型,发现当前模型在女性健康基准上的平均失败率约为60%,且不同专科和错误类型间表现差异显著。特别地,所有模型普遍在识别‘遗漏紧急情况’方面表现不佳;而较新模型如GPT-5在避免不当建议方面有明显提升。研究结果表明,现有AI聊天机器人尚无法可靠提供女性健康建议。

原文摘要 · Abstract (English)

As large language models (LLMs) become primary sources of health information for millions, their accuracy in women's health remains critically unexamined. We introduce the Women's Health Benchmark (WHB), the first benchmark evaluating LLM performance specifically in women's health. Our benchmark comprises 96 rigorously validated model stumps covering five medical specialties (obstetrics and gynecology, emergency medicine, primary care, oncology, and neurology), three query types (patient query, clinician query, and evidence/policy query), and eight error types (dosage/medication errors, missing critical information, outdated guidelines/treatment recommendations, incorrect treatment advice, incorrect factual information, missing/incorrect differential diagnosis, missed urgency, and inappropriate recommendations). We evaluated 13 state-of-the-art LLMs and revealed alarming gaps: current models show approximately 60\% failure rates on the women's health benchmark, with performance varying dramatically across specialties and error types. Notably, models universally struggle with "missed urgency" indicators, while newer models like GPT-5 show significant improvements in avoiding inappropriate recommendations. Our findings underscore that AI chatbots are not yet fully able of providing reliable advice in women's health.

大模型评测女性健康医疗AI可靠性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。