用医生定制的评分标准评估医疗AI,让大模型评分成本降千倍。
Case-Specific Rubrics for Clinical AI Evaluation: Methodology, Validation, and LLM-Clinician Agreement Across 823 Encounters

- 为每例临床场景设计医生撰写的评分细则,确保评估贴近真实诊疗。
- 医生评分与大模型评分一致性达0.42-0.46,优于医生间一致率。
- 大模型生成评分标准成本仅为人工的千分之一,适合大规模评估。
临床AI系统需具备临床有效、经济可行且能响应迭代更新的评估方法。逐例依赖专家评审的方法过于缓慢昂贵,难以支持安全迭代部署。本文提出一种基于病例特异性的医生撰写评分标准方法,并检验大语言模型(LLM)生成的评分标准能否接近医生共识。20位医生为823个临床案例(736个真实、87个合成)创建了1,646份评分标准,涵盖初级保健、精神科、肿瘤学及行为健康领域。每份评分标准通过验证:基于LLM的评分代理能稳定区分出医生偏好的输出优于被拒绝的输出。对七个嵌入电子病历系统的AI版本在全部案例中进行评估。结果显示,医生制定的评分标准能有效区分高质量与低质量输出(中位分数差距82.9%),评分稳定性高(中位范围0.00%)。中位得分从84%提升至95%。后续实验中,医生与大模型的排序一致性(tau: 0.42–0.46)达到或超过医生间一致性(tau: 0.38–0.43),归因于天花板压缩效应和大模型评分标准的改进。讨论表明,该收敛支持将大模型评分标准与医生标准并行使用。大模型评分标准成本仅约为人工的千分之一,显著提升评估覆盖范围,而持续的临床作者参与保障评估的专家判断基础。天花板压缩对未来的评价者间一致性研究构成方法论挑战。结论:病例特异性评分标准为临床AI评估提供了一条兼顾专家判断与自动化效率的路径,实现三数量级的成本降低。
原文摘要 · Abstract (English)
Objective. Clinical AI documentation systems require evaluation methodologies that are clinically valid, economically viable, and sensitive to iterative changes. Methods requiring expert review per scoring instance are too slow and expensive for safe, iterative deployment. We present a case-specific, clinician-authored rubric methodology for clinical AI evaluation and examine whether LLM-generated rubrics can approximate clinician agreement. Materials and Methods. Twenty clinicians authored 1,646 rubrics for 823 clinical cases (736 real-world, 87 synthetic) across primary care, psychiatry, oncology, and behavioral health. Each rubric was validated by confirming that an LLM-based scoring agent consistently scored clinician-preferred outputs higher than rejected ones. Seven versions of an EHR-embedded AI agent for clinicians were evaluated across all cases. Results. Clinician-authored rubrics discriminated effectively between high- and low-quality outputs (median score gap: 82.9%) with high scoring stability (median range: 0.00%). Median scores improved from 84% to 95%. In later experiments, clinician-LLM ranking agreement (tau: 0.42-0.46) matched or exceeded clinician-clinician agreement (tau: 0.38-0.43), attributable to both ceiling compression and LLM rubric improvement. Discussion. This convergence supports incorporating LLM rubrics alongside clinician-authored ones. At roughly 1,000 times lower cost, LLM rubrics enable substantially greater evaluation coverage, while continued clinical authorship grounds evaluation in expert judgment. Ceiling compression poses a methodological challenge for future inter-rater agreement studies. Conclusion. Case-specific rubrics offer a path for clinical AI evaluation that preserves expert judgment while enabling automation at three orders lower cost. Clinician-authored rubrics establish the baseline against which LLM rubrics are validated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。