arXiv:2510.03700cs.AI2025-10

提出分层评估框架H-DDx,更准确衡量大模型在鉴别诊断中的临床价值。

H-DDx: A Hierarchical Evaluation Framework for Differential Diagnosis

  • 构建检索重排管道将自由文本诊断映射到ICD-10编码
  • 采用分层指标奖励与真实诊断相近的预测结果
  • 揭示模型在整体临床上下文判断上的优势,适合医疗AI评估

准确的鉴别诊断对患者治疗和预后至关重要。近期大型语言模型(LLMs)被用于从患者主诉中生成鉴别诊断列表,但现有评估多依赖于平铺直叙的指标(如Top-k准确率),无法区分临床相关的近似误诊与远距离错误。为此,我们提出H-DDx——一种分层评估框架,通过检索与重排流程将自由文本诊断映射至ICD-10编码,并采用分层度量标准,对与真实诊断相近的预测给予奖励。在22个领先模型的基准测试中,传统平铺指标低估了性能,忽略了具有临床意义的输出;我们的结果凸显了领域专用开源模型的优势。此外,该框架提升了可解释性,揭示出即使精确诊断失败,模型仍常能正确识别更广泛的临床背景。

原文摘要 · Abstract (English)

An accurate differential diagnosis (DDx) is essential for patient care, shaping therapeutic decisions and influencing outcomes. Recently, Large Language Models (LLMs) have emerged as promising tools to support this process by generating a DDx list from patient narratives. However, existing evaluations of LLMs in this domain primarily rely on flat metrics, such as Top-k accuracy, which fail to distinguish between clinically relevant near-misses and diagnostically distant errors. To mitigate this limitation, we introduce H-DDx, a hierarchical evaluation framework that better reflects clinical relevance. H-DDx leverages a retrieval and reranking pipeline to map free-text diagnoses to ICD-10 codes and applies a hierarchical metric that credits predictions closely related to the ground-truth diagnosis. In benchmarking 22 leading models, we show that conventional flat metrics underestimate performance by overlooking clinically meaningful outputs, with our results highlighting the strengths of domain-specialized open-source models. Furthermore, our framework enhances interpretability by revealing hierarchical error patterns, demonstrating that LLMs often correctly identify the broader clinical context even when the precise diagnosis is missed.

鉴别诊断大模型评估ICD-10医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。