arXiv:2603.11413cs.HCcs.AI2026-03被引 4

测试发现,评估方式比模型能力更影响健康AI误判率。

Evaluation format, not model capability, drives triage failure in the assessment of consumer health AI

  • 用真实患者对话模式测试,准确率比考试式提问高6.4个百分点
  • 强制选A/B/C/D导致三款模型误判率从0%飙升至24%
  • 糖尿病酮症酸中毒可100%正确识别,哮喘识别率从48%升至80%

Ramaswamy等在《自然·医学》中报告,ChatGPT Health在51.6%的紧急情况中低估风险,得出消费者健康AI存在安全风险的结论。但其评估采用考试式协议——强制A/B/C/D选择、知识压制、禁止追问,与真实使用场景差异显著。我们对五款前沿大模型(GPT-5.2、Claude Sonnet 4.6、Claude Opus 4.6、Gemini 3 Flash、Gemini 3.1 Pro)在17个场景的复现数据集上进行测试,分别在受限(考试式,1,275次试验)和自然对话(患者式消息,850次试验)条件下开展,结合目标消融实验和提示忠实性验证。结果显示,自然交互使分诊准确率提升6.4个百分点(p=0.015)。糖尿病酮症酸中毒在所有模型与条件下均100%正确分诊;哮喘分诊率从48%提升至80%。强制选择是主要失败机制:三款模型在强制选择下得分0–24%,自由文本下达100%(全部p<10⁻⁸),其自主语言仍建议紧急就医,但强制格式记录为低估。对作者公开提示的忠实性检查确认,该框架产生依赖模型、依赖案例的结果。结果表明,标题中的低估率高度依赖评估格式,可能无法稳定反映实际部署表现。消费者健康AI的有效评估需在真实使用条件下进行。

原文摘要 · Abstract (English)

Ramaswamy et al. reported in Nature Medicine that ChatGPT Health under-triages 51.6% of emergencies, concluding that consumer-facing AI triage poses safety risks. However, their evaluation used an exam-style protocol -- forced A/B/C/D output, knowledge suppression, and suppression of clarifying questions -- that differs fundamentally from how consumers use health chatbots. We tested five frontier LLMs (GPT-5.2, Claude Sonnet 4.6, Claude Opus 4.6, Gemini 3 Flash, Gemini 3.1 Pro) on a 17-scenario partial replication bank under constrained (exam-style, 1,275 trials) and naturalistic (patient-style messages, 850 trials) conditions, with targeted ablations and prompt-faithful checks using the authors' released prompts. Naturalistic interaction improved triage accuracy by 6.4 percentage points ($p = 0.015$). Diabetic ketoacidosis was correctly triaged in 100% of trials across all models and conditions. Asthma triage improved from 48% to 80%. The forced A/B/C/D format was the dominant failure mechanism: three models scored 0--24% with forced choice but 100% with free text (all $p < 10^{-8}$), consistently recommending emergency care in their own words while the forced-choice format registered under-triage. Prompt-faithful checks on the authors' exact released prompts confirmed the scaffold produces model-dependent, case-dependent results. Our results suggest that the headline under-triage rate is highly contingent on evaluation format and may not generalize as a stable estimate of deployed triage behavior. Valid evaluation of consumer health AI requires testing under conditions that reflect actual use.

健康AI评估方法大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。