arXiv:2509.10584cs.CYcs.AI2025-09被引 1

用大模型分析社交媒体,自动识别符合临床试验条件的潜在参与者。

Smart Trial: Evaluating the Use of Large Language Models for Recruiting Clinical Trial Participants via Social Media

  • 基于社交媒体内容,用大模型判断用户是否满足试验条件。
  • 在结直肠癌和前列腺癌社群中,模型准确率尚不理想,需多步推理。
  • 适合医疗数据挖掘、AI辅助临床研究者参考。

临床试验是推动医学进步的关键,但高效招募符合条件的参与者仍面临挑战。传统方式如广告或医院电子病历筛查耗时且受限于地理范围。本文利用社交媒体上用户分享的健康信息,探索大语言模型(LLMs)在识别潜在受试者方面的潜力。为此,构建了TRIALQA数据集,涵盖来自结直肠癌和前列腺癌子论坛的社交媒体文本,并由专业标注员标注用户是否符合真实临床试验的入组标准及其参与意愿原因。在两个任务上对七种主流LLM进行评估,采用六种训练与推理策略。实验表明,尽管大模型展现出潜力,但在处理复杂、多跳推理以准确判断入组资格方面仍存在显著困难。

原文摘要 · Abstract (English)

Clinical trials (CT) are essential for advancing medical research and treatment, yet efficiently recruiting eligible participants -- each of whom must meet complex eligibility criteria -- remains a significant challenge. Traditional recruitment approaches, such as advertisements or electronic health record screening within hospitals, are often time-consuming and geographically constrained. This work addresses the recruitment challenge by leveraging the vast amount of health-related information individuals share on social media platforms. With the emergence of powerful large language models (LLMs) capable of sophisticated text understanding, we pose the central research question: Can LLM-driven tools facilitate CT recruitment by identifying potential participants through their engagement on social media? To investigate this question, we introduce TRIALQA, a novel dataset comprising two social media collections from the subreddits on colon cancer and prostate cancer. Using eligibility criteria from public real-world CTs, experienced annotators are hired to annotate TRIALQA to indicate (1) whether a social media user meets a given eligibility criterion and (2) the user's stated reasons for interest in participating in CT. We benchmark seven widely used LLMs on these two prediction tasks, employing six distinct training and inference strategies. Our extensive experiments reveal that, while LLMs show considerable promise, they still face challenges in performing the complex, multi-hop reasoning needed to accurately assess eligibility criteria.

临床试验大模型应用社交媒体医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。