为LLM从病历中提取数据的准确性提供验证框架
Ensuring Reliability of Curated EHR-Derived Data: The Validation of Accuracy for LLM/ML-Extracted Information and Data (VALID) Framework
- 构建多维度评估体系,对比专家标注与LLM提取结果
- 通过一致性检查和重复分析发现潜在错误
- 支持按人群分组评估偏差,适合医学AI研究者使用
大型语言模型(LLMs)在从电子健康记录(EHRs)中提取临床数据方面应用日益广泛,显著提升了肿瘤学真实世界数据(RWD)整理的可扩展性和效率。然而,其使用带来了可靠性、准确性和公平性方面的挑战,这些对研究、监管和临床应用至关重要。现有真实世界数据与人工智能的质量保障框架未能充分应对LLM提取数据特有的错误模式与复杂性。本文提出一个全面的临床数据提取质量评估框架,整合变量级性能基准测试(对比专家人工提取)、自动化内部一致性与合理性验证,以及与人工标注数据集或外部标准的重复分析。该多维度方法可识别需改进的变量,系统检测隐性错误,并确认数据集在真实世界研究中的适用性。同时,框架支持按人口统计学子组分层评估偏差。该方法提供了严谨透明的评估手段,推动行业标准进步,支持肿瘤学研究与实践中可信AI证据生成。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to extract clinical data from electronic health records (EHRs), offering significant improvements in scalability and efficiency for real-world data (RWD) curation in oncology. However, the adoption of LLMs introduces new challenges in ensuring the reliability, accuracy, and fairness of extracted data, which are essential for research, regulatory, and clinical applications. Existing quality assurance frameworks for RWD and artificial intelligence do not fully address the unique error modes and complexities associated with LLM-extracted data. In this paper, we propose a comprehensive framework for evaluating the quality of clinical data extracted by LLMs. The framework integrates variable-level performance benchmarking against expert human abstraction, automated verification checks for internal consistency and plausibility, and replication analyses comparing LLM-extracted data to human-abstracted datasets or external standards. This multidimensional approach enables the identification of variables most in need of improvement, systematic detection of latent errors, and confirmation of dataset fitness-for-purpose in real-world research. Additionally, the framework supports bias assessment by stratifying metrics across demographic subgroups. By providing a rigorous and transparent method for assessing LLM-extracted RWD, this framework advances industry standards and supports the trustworthy use of AI-powered evidence generation in oncology research and practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。