用大模型提升临床试验匹配效率,降低人工错误与招募延迟。
A systematic review of trial-matching pipelines using large language models
- 基于大模型构建多任务匹配管道,支持患者与试验条款自动比对。
- GPT-4在匹配与筛选准确率上优于其他模型,但成本更高。
- 适合医疗数据隐私要求高、需轻量化部署的医院或研究机构。
将患者匹配至临床试验是发现新疗法的关键,尤其在肿瘤学领域。然而,人工匹配耗时且易出错,导致招募延迟。结合大语言模型(LLMs)的匹配管道提供了有前景的解决方案。本研究系统回顾了2020至2025年间从三个学术数据库和一个预印本平台收录的文献,共筛选出126篇相关文章,其中31篇符合纳入标准。研究聚焦于患者-条目匹配(n=4)、患者-试验匹配(n=10)、试验-患者匹配(n=2)、二分类资格判断(n=1)或综合任务(n=14)。16项使用合成数据,14项使用真实患者数据,1项同时使用。数据集与评估指标差异导致跨研究可比性差。在直接对比中,GPT-4模型在匹配与资格提取方面持续优于其他模型,包括微调版本,尽管成本更高。有效策略包括零样本提示、专用模型如GPT-4o、先进检索方法,以及微调小型开源模型以保障数据隐私。关键挑战包括获取足够大的真实世界数据集,以及部署中的成本控制、幻觉、数据泄露与偏见风险。该综述总结了大模型在临床试验匹配中的进展,指明未来方向与主要局限。标准化指标、更真实的测试集,以及对成本效益与公平性的关注,对广泛部署至关重要。
原文摘要 · Abstract (English)
Matching patients to clinical trial options is critical for identifying novel treatments, especially in oncology. However, manual matching is labor-intensive and error-prone, leading to recruitment delays. Pipelines incorporating large language models (LLMs) offer a promising solution. We conducted a systematic review of studies published between 2020 and 2025 from three academic databases and one preprint server, identifying LLM-based approaches to clinical trial matching. Of 126 unique articles, 31 met inclusion criteria. Reviewed studies focused on matching patient-to-criterion only (n=4), patient-to-trial only (n=10), trial-to-patient only (n=2), binary eligibility classification only (n=1) or combined tasks (n=14). Sixteen used synthetic data; fourteen used real patient data; one used both. Variability in datasets and evaluation metrics limited cross-study comparability. In studies with direct comparisons, the GPT-4 model consistently outperformed other models, even finely-tuned ones, in matching and eligibility extraction, albeit at higher cost. Promising strategies included zero-shot prompting with proprietary LLMs like the GPT-4o model, advanced retrieval methods, and fine-tuning smaller, open-source models for data privacy when incorporation of large models into hospital infrastructure is infeasible. Key challenges include accessing sufficiently large real-world data sets, and deployment-associated challenges such as reducing cost, mitigating risk of hallucinations, data leakage, and bias. This review synthesizes progress in applying LLMs to clinical trial matching, highlighting promising directions and key limitations. Standardized metrics, more realistic test sets, and attention to cost-efficiency and fairness will be critical for broader deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。