arXiv:2503.15374cs.CLcs.AI2025-03被引 9

用多模态大模型自动匹配临床试验患者,提升效率与准确率。

Real-world validation of a multimodal LLM-powered pipeline for High-Accuracy Clinical Trial Patient Matching leveraging EHR data

  • 基于大模型的多模态管道,无需定制集成即可处理电子病历
  • 在真实数据上达到87%匹配准确率,单人审核时间缩短80%
  • 适合希望快速部署AI患者匹配系统的医疗机构

背景:临床试验患者招募受限于复杂的入选标准和耗时的病历审查。以往仅依赖文本的模型因推理能力有限、图像转文字导致信息损失、缺乏通用病历系统集成而难以可靠扩展。方法:提出一种无需集成的通用型大模型流程,直接处理电子病历原始文档。利用最新大模型的推理能力评估复杂标准,通过多模态能力解析医疗记录避免信息丢失,并采用多模态嵌入实现高效病历检索。在n2c2 2018队列选择数据集(288名糖尿病患者)和由30个机构485名患者组成的现实世界数据集(匹配36项不同试验)上验证。结果:在n2c2数据集上达到93%的准则级准确率新高;在真实世界中准确率为87%,主要受病历信息不足影响。用户平均每人每例审核时间低于9分钟,较传统人工方式提升80%。结论:该流程无需定制化适配即可实现跨机构规模化部署,显著提升临床试验患者匹配效率。

原文摘要 · Abstract (English)

Background: Patient recruitment in clinical trials is hindered by complex eligibility criteria and labor-intensive chart reviews. Prior research using text-only models have struggled to address this problem in a reliable and scalable way due to (1) limited reasoning capabilities, (2) information loss from converting visual records to text, and (3) lack of a generic EHR integration to extract patient data. Methods: We introduce a broadly applicable, integration-free, LLM-powered pipeline that automates patient-trial matching using unprocessed documents extracted from EHRs. Our approach leverages (1) the new reasoning-LLM paradigm, enabling the assessment of even the most complex criteria, (2) visual capabilities of latest LLMs to interpret medical records without lossy image-to-text conversions, and (3) multimodal embeddings for efficient medical record search. The pipeline was validated on the n2c2 2018 cohort selection dataset (288 diabetic patients) and a real-world dataset composed of 485 patients from 30 different sites matched against 36 diverse trials. Results: On the n2c2 dataset, our method achieved a new state-of-the-art criterion-level accuracy of 93\%. In real-world trials, the pipeline yielded an accuracy of 87\%, undermined by the difficulty to replicate human decision-making when medical records lack sufficient information. Nevertheless, users were able to review overall eligibility in under 9 minutes per patient on average, representing an 80\% improvement over traditional manual chart reviews. Conclusion: This pipeline demonstrates robust performance in clinical trial patient matching without requiring custom integration with site systems or trial-specific tailoring, thereby enabling scalable deployment across sites seeking to leverage AI for patient matching.

临床试验大模型电子病历患者匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。