arXiv:2607.18828cs.AI2026-07

测试医疗AI在信息缺失下的安全表现,发现评判者差异会显著影响结果

Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety

论文配图:Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety
图 1 · 摘自论文原文
  • 通过删除对话后半部分模拟信息缺失,用多模型评委评估AI响应安全性
  • 不同评委对AI安全性的评分差异大,同一模型在不同评委下排名可能反转
  • 大模型评委比临床医生更宽容,说明AI安全评估需考虑评判标准偏差

医疗AI的就绪性压力测试传统上聚焦封闭式与多模态基准。本文将其扩展至信息缺失下的开放对话场景,安全行为体现为识别缺失信息并合理回应,而非过度承诺;此时评估者本身也成为测量的一部分。我们通过删除HealthBench对话中用户最后一轮的后半部分,对四款模型(Claude Opus 4.8、GPT-5.5、Grok 4.3、Gemini 3.5 Flash)进行压力测试,采用四名提供者级LLM评委组和盲法临床专家锚定参考进行评分。结果显示:第一,评委选择显著影响评估结果——评委间一致性仅中等(Fleiss' kappa = 0.65),调整各评委普遍宽松度后,仍存在显著的同评委关联性(精确置换检验p=0.04;GPT-5.5相对概率提升+0.10),足以导致某模型在排除自身评委后排名变化;第二,大模型评委在50项盲评子集上普遍比临床医生更宽容——所有模型均更认可不确定性(66%-84%正确率对比52%),其中三个显著优于作者影响下的共识判断(判别一致性kappa=0.20-0.43)。在临床信息不足子集上,宽容差距扩大,模型排序保持一致。封闭式MedQA锚定验证显示准确率高,选项顺序效应在±5分区间内,说明安全差异源于校准问题而非知识缺陷。研究发布完整评测框架、提示、输出、评委组、扰动审计及人工标注协议。

原文摘要 · Abstract (English)

Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks. We extend it to open-ended clinical conversation under missing information, where safe behavior means recognizing absent information and qualifying, clarifying, or not over-committing - and where the evaluator becomes part of the measurement. We stress-test four models - three flagships (Claude Opus 4.8, GPT-5.5, Grok 4.3) and one mid-tier model (Gemini 3.5 Flash) - by deleting the latter half of the final user turn in HealthBench conversations, grading responses with a four-provider LLM-judge panel and a blinded clinician-anchored reference. Two evaluator-facing results are robust. First, judge choice materially changes apparent safety: inter-judge agreement is only moderate (Fleiss' kappa = 0.65), and after adjusting for each judge's general leniency (vote-level logistic regression), a positive same-provider association remains (exact permutation p = 0.04; GPT-5.5 ~ +0.10 on the probability scale) - large enough to change which model appears to over-commit least once its own-provider judge is excluded. Second, LLM judges are more permissive than clinicians on a blinded 50-item subsample: all four are significantly more lenient than the stricter independent clinician (crediting appropriate uncertainty on 66-84% of items vs 52%), and three of four than the author-influenced consensus (Grok directional only; judge-vs-consensus kappa = 0.20-0.43). On the author-audited clinical-underdetermined subset the permissiveness gap widened and the point-estimate model ordering held. A closed-ended MedQA anchor confirms accuracy is high and option-order effects are within a +/-5-point equivalence region for three of four models, so the safety gap is about calibration, not knowledge. We release the harness, prompts, per-item outputs, judge panel, perturbation audit, and human-annotation protocol.

医疗AI评估基准模型校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。