用多阶段框架让大模型可靠提取海量病历中的成瘾诊断信息
A Multi-Stage Validation Framework for Trustworthy Large-scale Clinical Information Extraction using Large Language Models
- 分六步验证:提示校准、规则过滤、语义验证、高阶模型复核等
- 91万份病历中提取出11类成瘾诊断,准确率F1达0.80
- 无需人工标注即可评估可信度,适合医疗数据大规模应用
大型语言模型(LLMs)在从非结构化医疗记录中提取临床信息方面展现出潜力,但其在真实场景中的应用受限于缺乏可扩展且可信的验证方法。传统评估依赖高成本的人工标注或不完整的结构化数据,难以实现群体规模应用。本文提出一种多阶段验证框架,支持在弱监督下对基于LLM的临床信息抽取进行严格评估。该框架整合提示校准、基于规则的合理性过滤、语义根基性评估、使用独立高容量判别模型的针对性确认评估、选择性专家审查及外部预测有效性分析,可在无需全面人工标注的情况下量化不确定性并刻画错误模式。我们将该框架应用于从919,783份临床笔记中提取11类物质使用障碍(SUD)诊断。规则过滤与语义根基性评估剔除了14.59%无依据、无关或结构上不合理的内容。对于高不确定案例,判别模型与领域专家评审结果高度一致(Gwet's AC1=0.80)。以判别模型输出为参考,主模型在宽松匹配标准下取得F1=0.80。LLM提取的SUD诊断还能更准确预测后续接受专门治疗(AUC=0.80),优于结构化数据基线。结果表明,无需密集标注即可实现可扩展、可信的LLM临床信息提取。
原文摘要 · Abstract (English)
Large language models (LLMs) show promise for extracting clinically meaningful information from unstructured health records, yet their translation into real-world settings is constrained by the lack of scalable and trustworthy validation approaches. Conventional evaluation methods rely heavily on annotation-intensive reference standards or incomplete structured data, limiting feasibility at population scale. We propose a multi-stage validation framework for LLM-based clinical information extraction that enables rigorous assessment under weak supervision. The framework integrates prompt calibration, rule-based plausibility filtering, semantic grounding assessment, targeted confirmatory evaluation using an independent higher-capacity judge LLM, selective expert review, and external predictive validity analysis to quantify uncertainty and characterize error modes without exhaustive manual annotation. We applied this framework to extraction of substance use disorder (SUD) diagnoses across 11 substance categories from 919,783 clinical notes. Rule-based filtering and semantic grounding removed 14.59% of LLM-positive extractions that were unsupported, irrelevant, or structurally implausible. For high-uncertainty cases, the judge LLM's assessments showed substantial agreement with subject matter expert review (Gwet's AC1=0.80). Using judge-evaluated outputs as references, the primary LLM achieved an F1 score of 0.80 under relaxed matching criteria. LLM-extracted SUD diagnoses also predicted subsequent engagement in SUD specialty care more accurately than structured-data baselines (AUC=0.80). These findings demonstrate that scalable, trustworthy deployment of LLM-based clinical information extraction is feasible without annotation-intensive evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。