解析医生在医疗AI评估中的分歧根源,发现信息缺失是主要可改进点。
Decomposing Physician Disagreement in HealthBench
- 通过分解标签差异,识别出医生个体和评分标准对分歧影响有限
- 81.8%的分歧来自病例本身,现有元数据无法降低该差异
- 可消除的不确定性(如信息缺失)使分歧概率翻倍,提示评估设计优化空间
我们对HealthBench医疗AI评估数据集中医生分歧进行分解,以理解其变异来源及可观测特征的解释力。评分标准身份解释了15.8%的标签变异,但仅占3.6-6.9%的分歧变异;医生身份仅贡献2.4%。主导的81.8%病例级残差未被HealthBench元数据标签(z = -0.22, p = 0.83)、规范评分语言(伪R² = 1.2%)、医学专科(0/300个Tukey对显著)、表面特征分诊(AUC = 0.58)或嵌入表示(AUC = 0.485)所缓解。分歧随完成质量呈倒U型分布(AUC = 0.689),表明医生对明确优劣结果共识高,对边界案例分歧大。经医生验证的不确定性分类显示,可减少的不确定性(如信息缺失、表述模糊)使分歧几率提升2.55倍(p < 10⁻²⁴),而不可减少的不确定性(真实医学歧义)无影响(OR = 1.01, p = 0.90),尽管前者仅解释约3%总变异。因此,医疗AI评估的共识上限具有结构性,但可减少与不可减少不确定性的区分表明,填补评估场景的信息空白可在非本质歧义处降低分歧,为评估设计提供可操作改进方向。
原文摘要 · Abstract (English)
We decompose physician disagreement in the HealthBench medical AI evaluation dataset to understand where variance resides and what observable features can explain it. Rubric identity accounts for 15.8% of met/not-met label variance but only 3.6-6.9% of disagreement variance; physician identity accounts for just 2.4%. The dominant 81.8% case-level residual is not reduced by HealthBench's metadata labels (z = -0.22, p = 0.83), normative rubric language (pseudo R^2 = 1.2%), medical specialty (0/300 Tukey pairs significant), surface-feature triage (AUC = 0.58), or embeddings (AUC = 0.485). Disagreement follows an inverted-U with completion quality (AUC = 0.689), confirming physicians agree on clearly good or bad outputs but split on borderline cases. Physician-validated uncertainty categories reveal that reducible uncertainty (missing context, ambiguous phrasing) more than doubles disagreement odds (OR = 2.55, p < 10^(-24)), while irreducible uncertainty (genuine medical ambiguity) has no effect (OR = 1.01, p = 0.90), though even the former explains only ~3% of total variance. The agreement ceiling in medical AI evaluation is thus largely structural, but the reducible/irreducible dissociation suggests that closing information gaps in evaluation scenarios could lower disagreement where inherent clinical ambiguity does not, pointing toward actionable evaluation design improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。