用多模态大模型自动匹配临床试验患者,提升效率与准确率。
Real-world validation of a multimodal LLM-powered pipeline for High-Accuracy Clinical Trial Patient Matching leveraging EHR data
- 基于大模型的多模态管道,无需定制集成即可处理电子病历
- 在真实数据上达到87%匹配准确率,单人审核时间缩短80%
- 适合希望快速部署AI患者匹配系统的医疗机构
背景:临床试验患者招募受限于复杂的入选标准和耗时的病历审查。以往仅依赖文本的模型因推理能力有限、图像转文字导致信息损失、缺乏通用病历系统集成而难以可靠扩展。方法:提出一种无需集成的通用型大模型流程,直接处理电子病历原始文档。利用最新大模型的推理能力评估复杂标准,通过多模态能力解析医疗记录避免信息丢失,并采用多模态嵌入实现高效病历检索。在n2c2 2018队列选择数据集(288名糖尿病患者)和由30个机构485名患者组成的现实世界数据集(匹配36项不同试验)上验证。结果:在n2c2数据集上达到93%的准则级准确率新高;在真实世界中准确率为87%,主要受病历信息不足影响。用户平均每人每例审核时间低于9分钟,较传统人工方式提升80%。结论:该流程无需定制化适配即可实现跨机构规模化部署,显著提升临床试验患者匹配效率。
原文摘要 · Abstract (English)
Background: Patient recruitment in clinical trials is hindered by complex eligibility criteria and labor-intensive chart reviews. Prior research using text-only models have struggled to address this problem in a reliable and scalable way due to (1) limited reasoning capabilities, (2) information loss from converting visual records to text, and (3) lack of a generic EHR integration to extract patient data. Methods: We introduce a broadly applicable, integration-free, LLM-powered pipeline that automates patient-trial matching using unprocessed documents extracted from EHRs. Our approach leverages (1) the new reasoning-LLM paradigm, enabling the assessment of even the most complex criteria, (2) visual capabilities of latest LLMs to interpret medical records without lossy image-to-text conversions, and (3) multimodal embeddings for efficient medical record search. The pipeline was validated on the n2c2 2018 cohort selection dataset (288 diabetic patients) and a real-world dataset composed of 485 patients from 30 different sites matched against 36 diverse trials. Results: On the n2c2 dataset, our method achieved a new state-of-the-art criterion-level accuracy of 93\%. In real-world trials, the pipeline yielded an accuracy of 87\%, undermined by the difficulty to replicate human decision-making when medical records lack sufficient information. Nevertheless, users were able to review overall eligibility in under 9 minutes per patient on average, representing an 80\% improvement over traditional manual chart reviews. Conclusion: This pipeline demonstrates robust performance in clinical trial patient matching without requiring custom integration with site systems or trial-specific tailoring, thereby enabling scalable deployment across sites seeking to leverage AI for patient matching.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。