arXiv:2604.00024cs.CLcs.AI2026-04

针对女性健康领域设计专业评估基准,揭示大模型临床失效风险。

WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics

论文配图:WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics
图 1 · 摘自论文原文
  • 由专家构建47个真实场景,覆盖10类女性健康问题。
  • 22个模型平均得分未超75%,最优仅达72.1%,多数存在安全漏洞。
  • 适合医疗AI研发与临床部署前的安全性评估使用。

大型语言模型在医疗指导中应用日益广泛,但女性健康领域在评测设计中仍被忽视。本文提出女性健康基准(WHBench),包含47个专家设计的临床场景,覆盖10个女性健康主题,旨在暴露过时指南、安全隐患、用药错误及公平性盲区等临床级失效模式。我们采用23项标准的评估体系,涵盖临床准确性、完整性、安全性、沟通质量、指令遵循、公平性、不确定性处理和指南一致性,引入安全加权惩罚与服务端分数重算机制。在3,102次响应尝试中(3,100次评分),无一模型平均分超过75%,最佳模型仅达72.1%。即使顶尖模型也存在低完全正确率与显著危害率波动。人工评估者间一致性在响应标签层面中等,但在模型排名上较高,支持WHBench用于模型比较评估,同时凸显临床部署需专家监督。WHBench为公开、故障导向的评测基准,助力更安全、更公平的女性健康AI发展。

原文摘要 · Abstract (English)

Large language models are increasingly used for medical guidance, but women's health remains under-evaluated in benchmark design. We present the Women's Health Benchmark (WHBench), a targeted evaluation suite of 47 expert-crafted scenarios across 10 women's health topics, designed to expose clinically meaningful failure modes including outdated guidelines, unsafe omissions, dosing errors, and equity-related blind spots. We evaluate 22 models using a 23-criterion rubric spanning clinical accuracy, completeness, safety, communication quality, instruction following, equity, uncertainty handling, and guideline adherence, with safety-weighted penalties and server-side score recalculation. Across 3,102 attempted responses (3,100 scored), no model mean performance exceeds 75 percent; the best model reaches 72.1 percent. Even top models show low fully correct rates and substantial variation in harm rates. Inter-rater reliability is moderate at the response label level but high for model ranking, supporting WHBench utility for comparative system evaluation while highlighting the need for expert oversight in clinical deployment. WHBench provides a public, failure-mode-aware benchmark to track safer and more equitable progress in womens health AI.

女性健康大模型评估医疗AI安全评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。