AI评阅结果受评分协议影响,有标准才更真实可信。
AI Rater Discrimination Depends on Scoring Protocol in Complex Clinical Decision-Making

- 用双评分协议对比大模型在糖尿病治疗决策中的评分表现
- 无标准评分使分数集中在74-78分,有标准时差异达7.69至49.64分
- 带标准的评分能更好区分不同模型输出,适合临床复杂决策评估
临床AI评估日益依赖大语言模型作为评阅者,但其在不同评估条件下的评分行为尚未被量化。本研究通过因子实验分析了成人2型糖尿病(T2D)药物治疗在12个月随访中的AI评阅行为,该任务涉及复杂决策,以七项评估问题形式呈现。四个开源LLM同时担任临床决策支持系统(CDSS)模型和AI评阅者。每个CDSS输出在两种评分协议下被评分:基于患者特异性评分标准的黄金评分协议(GR),以及无评分标准的非黄金评分协议(Non-GR)。线性混合效应模型将评分协议与五个设计因素(CDSS模型、提示配置[文档引用生成(DRG)vs.基线]、评阅模型、提示字符、提示类型)交叉分析,估计主效应及其协议交互作用。结果显示,在所有问题中,非黄金评分协议下评阅者得分集中于74–78分区间,而黄金评分协议下平均分低7.69至49.64分,四分位距范围宽1.68至3.67倍。在每项问题中,黄金评分协议使评阅者对DRG与基线输出的区分能力提升1.76至5.10倍,同时揭示出评阅模型间显著的行为差异,而非黄金评分协议则掩盖了这些差异。结果表明,评分标准锚定是保留临床AI评估判别力的关键;当问题需患者或地域特定标准时,仅靠参数知识无法推断的非标准评分不可替代。
原文摘要 · Abstract (English)
Clinical AI evaluation increasingly delegates scoring to large language models (LLMs) acting as AI raters, yet their scoring behavior across evaluation conditions has not been quantitatively characterized. We address this gap through a factorial study of AI rater behavior in adult type 2 diabetes (T2D) pharmacotherapy at 12-month outpatient follow-up, a clinical task involving complex decision-making operationalized across seven evaluation questions. Four open-source LLMs served simultaneously as clinical decision support system (CDSS) models and AI raters. Each CDSS output was scored under two scoring protocols: a rubric-anchored Gold Rubric (GR) protocol incorporating a patient-specific rubric, and a rubric-free Non Gold Rubric (Non-GR) protocol. Linear mixed effects models crossed the scoring protocol factor with five design factors -- CDSS model, CDSS prompt configuration (document-referenced generation [DRG] vs.\ Baseline), rater model, prompt character, and prompt type -- and estimated main effects together with their protocol interactions. Across all questions, AI raters yielded consistently higher scores within a very narrow range (74--78 points on average) under Non-GR compared to those under GR (7.69 to 49.64 points lower mean scores; 1.68 to 3.67 times wider interquartile ranges). Within each question, GR amplified the AI rater's discrimination between DRG and Baseline CDSS outputs by factors of 1.76 to 5.10, while also revealing substantial behavioral variation across rater models that Non-GR suppressed. These findings support rubric anchoring as the scoring protocol that preserves discriminative power in clinical AI evaluation; rubric-free scoring cannot substitute when questions require patient-specific or jurisdiction-specific criteria that rater models cannot infer from parametric knowledge alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。