提出新评估方法,让医疗AI评价更贴近真实专家判断。
How to Evaluate Medical AI
- 用多专家意见作参照,评估AI诊断性能。
- 98%准确率实现自由文本诊断识别,效果极佳。
- 发现专家间差异常大于人机差异,支持相对评估。
人工智能融入医学诊断流程需可靠评价方法以确保可靠性与临床相关性。传统精度与召回率忽略专家判断的固有变异,导致评估不一致;虽κ系数更稳定,但解释性差。本文提出相对精度与召回率(RPAD/RRAD),将AI输出与多位专家意见对比,通过归一化处理专家分歧,提供更稳定真实的诊断质量度量。研究使用360个医学对话,对比多个大语言模型(LLM)与医生团队表现。结果表明,顶级模型如DeepSeek-V3在一致性上可媲美甚至超过专家共识。更重要的是,我们开发了一种自动化自由文本诊断识别方法,在该框架下达到98%准确率。同时发现,专家间判断差异往往大于人机差异,凸显绝对指标局限性,支持采用相对评估方法。
原文摘要 · Abstract (English)
The integration of artificial intelligence (AI) into medical diagnostic workflows requires robust and consistent evaluation methods to ensure reliability, clinical relevance, and the inherent variability in expert judgments. Traditional metrics like precision and recall often fail to account for the inherent variability in expert judgments, leading to inconsistent assessments of AI performance. Inter-rater agreement statistics like Cohen's Kappa are more reliable but they lack interpretability. We introduce Relative Precision and Recall of Algorithmic Diagnostics (RPAD and RRAD) - a new evaluation metrics that compare AI outputs against multiple expert opinions rather than a single reference. By normalizing performance against inter-expert disagreement, these metrics provide a more stable and realistic measure of the quality of predicted diagnosis. In addition to the comprehensive analysis of diagnostic quality measures, our study contains a very important side result. Our evaluation methodology allows us to avoid selecting diagnoses from a limited list when evaluating a given case. Instead, both the models being tested and the examiners verifying them arrive at a free-form diagnosis. In this automated methodology for establishing the identity of free-form clinical diagnoses, a remarkable 98% accuracy becomes attainable. We evaluate our approach using 360 medical dialogues, comparing multiple large language models (LLMs) against a panel of physicians. Large-scale study shows that top-performing models, such as DeepSeek-V3, achieve consistency on par with or exceeding expert consensus. Moreover, we demonstrate that expert judgments exhibit significant variability - often greater than that between AI and humans. This finding underscores the limitations of any absolute metrics and supports the need to adopt relative metrics in medical AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。