用大模型评估医疗文本效果,但存在盲点与偏见,需新框架保障安全。
A Scoping Review of LLM-as-a-Judge in Healthcare and the MedJUDGE Framework
- 提出医疗领域大模型评分框架MedJUDGE,分层管控风险
- 49项研究中仅1项上线,36%无专家参与,偏见评估严重不足
- 适合关注医疗AI评估安全的开发者与监管者
随着大语言模型(LLMs)在临床文本生成与处理中的广泛应用,可扩展的评估机制日益重要。LLM-as-a-Judge(LaaJ)利用大模型自动评估输出,替代昂贵的人工评审,但其在医疗领域的应用引发安全与偏见担忧。我们基于PRISMA-ScR方法,系统检索六个数据库(2020年1月至2026年1月),筛选出11,727篇文献,最终纳入49项研究。其中75.5%为评估与基准测试应用,85.7%采用逐项打分方式,73.5%使用GPT系列模型作为评判者。尽管应用增长迅速,验证严谨性仍有限:36项含人类参与的研究中,专家平均仅3人;13项(26.5%)完全无专家介入。36项(73.5%)未进行偏见风险测试,仅1项(2.0%)分析了人口统计公平性,无人评估时间稳定性或患者上下文影响。部署进展缓慢,仅1项(2.0%)进入生产环境,4项(8.2%)处于原型阶段。值得注意的是,当评判模型与被评系统共享训练数据或架构时,可能继承相同盲点,且一致性指标无法区分真实有效性与共通错误。当前缺乏人类监督、偏见评估薄弱及模型同质化,构成治理缺口,可能导致关键临床错误被忽略。为此,我们提出MedJUDGE(Medical Judge Utility, De-biasing, Governance and Evaluation)框架,基于临床风险层级,从有效性、安全性与问责性三方面构建分层评估体系,为医疗LaaJ系统提供部署导向的评估指导。
原文摘要 · Abstract (English)
As large language models (LLMs) increasingly generate and process clinical text, scalable evaluation has become critical. LLM-as-a-Judge (LaaJ), which uses LLMs to evaluate model outputs, offers a scalable alternative to costly expert review, but its healthcare adoption raises safety and bias concerns. We conducted a PRISMA-ScR scoping review of six databases (January 2020-January 2026), screening 11,727 studies and including 49. The landscape was dominated by evaluation and benchmarking applications (n=37, 75.5%), pointwise scoring (n=42, 85.7%), and GPT-family judges (n=36, 73.5%). Despite growing adoption, validation rigor was limited: among 36 studies with human involvement, the median number of expert validators was 3, while 13 (26.5%) used none. Risk of bias testing was absent in 36 studies (73.5%), only 1 (2.0%) examined demographic fairness, and none assessed temporal stability or patient context. Deployment remained limited, with 1 study (2.0%) reaching production and four (8.2%) prototype stage. Importantly, these gaps may interact: when judges and evaluated systems share training data or architectures, they may inherit similar blind spots, and agreement metrics may fail to distinguish true validity from shared errors. Minimal human oversight, limited bias assessment, and model monoculture together represent a governance gap where current validation may miss clinically significant errors. To address this, we propose MedJUDGE (Medical Judge Utility, De-biasing, Governance and Evaluation), a risk-stratified three-pillar framework organized around validity, safety, and accountability across clinical risk tiers, providing deployment-oriented evaluation guidance for healthcare LaaJ systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。