arXiv:2601.06300cs.CLcs.AI2026-01被引 3

预测临床试验入选标准变更,提升研究设计效率

$\texttt{AMEND++}$: Benchmarking Eligibility Criteria Amendments in Clinical Trials

  • 提出新NLP任务:预测临床试验入选标准是否会修改
  • 构建AMEND++数据集,包含历史版本与变更标签
  • 创新预训练方法CAMLM,显著提升预测准确率

临床试验修订常导致延迟、成本增加和管理负担,其中入选标准是最常修改的部分。本文提出‘入选标准修订预测’这一新型NLP任务,旨在预测初始试验方案中的入选标准是否会未来发生修改。为此,我们发布了AMEND++基准套件,包含两个数据集:AMEND,收录公开临床试验的入选标准版本历史与修订标签;AMEND_LLM,通过基于大模型的去噪流程筛选出的精炼子集,用于分离实质性修改。我们进一步提出‘感知变更的掩码语言建模’(CAMLM),一种利用历史编辑信息进行修订敏感表示学习的预训练策略。在多种基线上的实验表明,CAMLM持续提升预测性能,有助于实现更稳健、低成本的临床试验设计。

原文摘要 · Abstract (English)

Clinical trial amendments frequently introduce delays, increased costs, and administrative burden, with eligibility criteria being the most commonly amended component. We introduce \textit{eligibility criteria amendment prediction}, a novel NLP task that aims to forecast whether the eligibility criteria of an initial trial protocol will undergo future amendments. To support this task, we release $\texttt{AMEND++}$, a benchmark suite comprising two datasets: $\texttt{AMEND}$, which captures eligibility-criteria version histories and amendment labels from public clinical trials, and $\verb|AMEND_LLM|$, a refined subset curated using an LLM-based denoising pipeline to isolate substantive changes. We further propose $\textit{Change-Aware Masked Language Modeling}$ (CAMLM), a revision-aware pretraining strategy that leverages historical edits to learn amendment-sensitive representations. Experiments across diverse baselines show that CAMLM consistently improves amendment prediction, enabling more robust and cost-effective clinical trial design.

临床试验NLP预测模型数据标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。