arXiv:2607.14385cs.CLcs.LG2026-07

首个针对孕产与儿童健康诊断的反事实鲁棒性评测基准,揭示大模型诊断稳定性短板。

MamaBench: Benchmarking LLM Robustness in Maternal and Child Health Diagnosis through Counterfactual Clinical Perturbation

论文配图:MamaBench: Benchmarking LLM Robustness in Maternal and Child Health Diagnosis through Counterfactual Clinical Perturbation
图 1 · 摘自论文原文
  • 构建434个临床案例对,通过反事实扰动测试模型是否能区分相似病症
  • 基线准确率虚高16-28个百分点,真实鲁棒性不足65%
  • 提出证据锚定检索增强生成,提升诊断一致性,适合医疗AI开发者参考

大型语言模型在医学基准测试中表现优异,但这些测试孤立评估每个问题,无法衡量系统区分临床相似症状(需不同干预)的能力。我们提出MamaBench,首个针对孕产与儿科人工智能的反事实评测基准:包含434个专家撰写的临床叙事,共217对案例,覆盖371种病理,通过偏见陷阱率(BTR)评估模型在基础任务正确时仍失败反事实任务的条件概率。我们提出证据锚定RAG(EA-RAG),采用三阶段检索方法,以临床参数提取、覆盖率审计和对比子查询替代传统相似度聚合,实现更精准的证据覆盖。在四种前沿大模型的八种配置下,基础准确率比鲁棒准确率高16-28个百分点。EA-RAG在Claude Sonnet 4.6上实现20.3% BTR和65.0%鲁棒准确率,相比基线降低5.5个百分点,且不损害基础性能。剩余20% BTR表明,临床AI的反事实鲁棒性仍是开放挑战。

原文摘要 · Abstract (English)

Large language models achieve strong scores on medical benchmarks, yet these benchmarks evaluate each question in isolation, providing no measure of whether a system can distinguish clinically similar presentations requiring different interventions. We introduce MamaBench, the first counterfactual benchmark for maternal and paediatric AI: 434 expert-authored clinical narratives in 217 pairs across 371 pathologies, evaluated via the Bias Trap Rate (BTR), the conditional probability that a model fails the counterfactual given success on the base case. We propose Evidence-Anchored RAG (EA-RAG), a three-stage retrieval method that replaces aggregate similarity with an evidence coverage objective through clinical parameter extraction, coverage auditing, and contrastive sub-queries. Across eight configurations of four frontier LLMs, base accuracy overstates robust accuracy by 16-28 percentage points in every model. EA-RAG achieves 20.3% BTR and 65.0% robust accuracy on Claude Sonnet 4.6, a 5.5 percentage point BTR reduction without degrading base accuracy. The residual 20% BTR confirms that counterfactual robustness in clinical AI remains an open challenge. Keywords: counterfactual evaluation, clinical AI, maternal healthcare, retrieval-augmented generation, diagnostic robustness

临床AI反事实评测大模型鲁棒性医疗诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。