用可解释的语义轴引导生成临床试验入组标准,提升效率与可读性。
POET: Protocol Optimization via Eligibility Tuning
- 引入人口、检验指标等语义轴引导生成,无需预设具体实体
- 在自动评估、评分框架和医生评价中均优于无引导生成
- 适合临床研究者使用,兼顾实用性与可解释性
入组标准(EC)是临床试验设计的关键,但撰写过程对临床医生而言耗时且认知负荷高。现有自动化方法或需高度结构化输入(如预定义实体),或依赖端到端系统从极少输入生成完整标准,实用性受限。本文提出一种引导式生成框架,利用大语言模型生成可解释的语义轴(如人口学特征、实验室参数、行为因素),作为生成指引,使临床医生可在不指定具体实体的情况下有效控制生成内容。同时,构建基于评分表的可复用评估框架,从临床有意义维度评估生成标准。实验表明,该引导式方法在自动评估、评分框架评估及临床医生评价中均显著优于无引导生成,为人工智能辅助试验设计提供了实用且可解释的解决方案。
原文摘要 · Abstract (English)
Eligibility criteria (EC) are essential for clinical trial design, yet drafting them remains a time-intensive and cognitively demanding task for clinicians. Existing automated approaches often fall at two extremes either requiring highly structured inputs, such as predefined entities to generate specific criteria, or relying on end-to-end systems that produce full eligibility criteria from minimal input such as trial descriptions limiting their practical utility. In this work, we propose a guided generation framework that introduces interpretable semantic axes, such as Demographics, Laboratory Parameters, and Behavioral Factors, to steer EC generation. These axes, derived using large language models, offer a middle ground between specificity and usability, enabling clinicians to guide generation without specifying exact entities. In addition, we present a reusable rubric-based evaluation framework that assesses generated criteria along clinically meaningful dimensions. Our results show that our guided generation approach consistently outperforms unguided generation in both automatic, rubric-based and clinician evaluations, offering a practical and interpretable solution for AI-assisted trial design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。