arXiv:2501.08977cs.AI2025-01被引 32

为大模型生成的病历摘要设计了可验证的质量评估工具。

Development and Validation of the Provider Documentation Summarization Quality Instrument for Large Language Models

  • 构建针对大模型摘要的多维度评估框架,涵盖结构与内容
  • 在779份真实病历摘要上验证,内部一致性系数达0.879
  • 适合临床医生和研究者用于评估医疗大模型输出质量

随着大型语言模型(LLMs)被集成到电子健康记录(EHR)工作流程中,需要经过验证的工具来评估其性能。现有医生文档质量评估工具通常不适用于大模型生成文本,且缺乏在真实数据上的验证。本文开发了医生文档摘要质量评估工具(PDSQI-9),用于评估大模型生成的临床摘要。使用GPT-4o、Mixtral 8x7b和Llama 3-8b等多个大模型,从多专科真实EHR数据生成多文档摘要。验证包括皮尔逊相关性(内容效度)、因子分析与克朗巴哈α(结构效度)、评分者间一致性(ICC与Krippendorff's alpha,泛化性)、半德尔菲法(内容效度)以及高质量与低质量摘要的对比(区分效度)。七名医生评估了779份摘要,回答8,329个问题,评分者间可靠性检验具有超过80%的统计功效。PDSQI-9表现出强内部一致性(克朗巴哈α = 0.879;95% CI: 0.867–0.891)和高评分者间一致性(ICC = 0.867;95% CI: 0.867–0.868),支持结构效度与泛化性。因子分析识别出解释58%方差的四因子模型:组织性、清晰性、准确性与实用性。内容效度由笔记长度与得分的相关性支持(简洁性:rho = -0.200, p = 0.029;组织性:ρ = -0.190, p = 0.037)。区分效度成功区分高低质量摘要(p < 0.001)。PDSQI-9具备稳健的构念效度,可用于临床实践,促进大模型在医疗工作流中的安全整合。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) are integrated into electronic health record (EHR) workflows, validated instruments are essential to evaluate their performance before implementation. Existing instruments for provider documentation quality are often unsuitable for the complexities of LLM-generated text and lack validation on real-world data. The Provider Documentation Summarization Quality Instrument (PDSQI-9) was developed to evaluate LLM-generated clinical summaries. Multi-document summaries were generated from real-world EHR data across multiple specialties using several LLMs (GPT-4o, Mixtral 8x7b, and Llama 3-8b). Validation included Pearson correlation for substantive validity, factor analysis and Cronbach's alpha for structural validity, inter-rater reliability (ICC and Krippendorff's alpha) for generalizability, a semi-Delphi process for content validity, and comparisons of high-versus low-quality summaries for discriminant validity. Seven physician raters evaluated 779 summaries and answered 8,329 questions, achieving over 80% power for inter-rater reliability. The PDSQI-9 demonstrated strong internal consistency (Cronbach's alpha = 0.879; 95% CI: 0.867-0.891) and high inter-rater reliability (ICC = 0.867; 95% CI: 0.867-0.868), supporting structural validity and generalizability. Factor analysis identified a 4-factor model explaining 58% of the variance, representing organization, clarity, accuracy, and utility. Substantive validity was supported by correlations between note length and scores for Succinct (rho = -0.200, p = 0.029) and Organized ($ρ= -0.190$, $p = 0.037$). Discriminant validity distinguished high- from low-quality summaries ($p < 0.001$). The PDSQI-9 demonstrates robust construct validity, supporting its use in clinical practice to evaluate LLM-generated summaries and facilitate safer integration of LLMs into healthcare workflows.

医疗AI评估工具大模型测评

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。