arXiv:2501.11114cs.CLcs.AI2025-01被引 3

用大模型自动筛选临床试验患者,提升效率但对精细推理仍有限。

Clinical trial cohort selection using Large Language Models on n2c2 Challenges

  • 基于n2c2挑战赛数据,测试大模型在患者筛选中的表现。
  • 简单筛选任务效果良好,但涉及细粒度知识时准确率下降。
  • 适合关注医疗文本自动化筛选的研究者或临床研发人员。

临床试验是医学领域引入新疗法和创新的关键环节。然而,患者队列筛选过程耗时,通常需人工查阅病历文本以查找特定关键词。尽管已有研究致力于统一各平台信息标准,自然语言处理(NLP)工具在识别文本报告中的入组标准方面仍至关重要。近年来,预训练大语言模型(LLMs)因具备对文本的深层理解能力,在多种NLP任务中受到广泛关注。本文研究了大语言模型在临床试验队列筛选任务中的表现,并利用n2c2挑战赛数据集进行基准测试。结果显示,大模型在简单队列筛选任务中表现令人鼓舞,但在需要精细知识和推理的任务上仍面临挑战。

原文摘要 · Abstract (English)

Clinical trials are a critical process in the medical field for introducing new treatments and innovations. However, cohort selection for clinical trials is a time-consuming process that often requires manual review of patient text records for specific keywords. Though there have been studies on standardizing the information across the various platforms, Natural Language Processing (NLP) tools remain crucial for spotting eligibility criteria in textual reports. Recently, pre-trained large language models (LLMs) have gained popularity for various NLP tasks due to their ability to acquire a nuanced understanding of text. In this paper, we study the performance of large language models on clinical trial cohort selection and leverage the n2c2 challenges to benchmark their performance. Our results are promising with regard to the incorporation of LLMs for simple cohort selection tasks, but also highlight the difficulties encountered by these models as soon as fine-grained knowledge and reasoning are required.

临床试验大模型NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。