用机器学习提前预测临床试验的用药错误风险,提升研究安全性。
Early Risk Stratification of Dosing Errors in Clinical Trials Using Machine Learning
- 融合结构化数据与文本信息,构建多模态预测模型
- 晚期融合模型AUC达0.862,可准确识别高风险试验
- 概率校准后实现可解释的风险分层,适合临床质控参考
本研究旨在开发一种基于机器学习(ML)的框架,利用试验启动前的可用信息,对临床试验(CTs)进行早期用药错误风险分层。从ClinicalTrials.gov构建包含42,112项试验的数据集,提取结构化、半结构化数据及协议相关的非结构化文本。通过不良事件报告、MedDRA术语和威尔逊置信区间,为试验打上二元标签(是否具高用药错误率)。评估了仅使用结构特征的XGBoost模型、基于文本的ClinicalModernBERT模型,以及结合两者的简单晚期融合模型。采用事后概率校准,实现可解释的试验级风险分层。结果显示,晚期融合模型表现最佳,AUC-ROC为0.862。校准后的输出能稳健地将试验划分为预设风险等级,高预测风险组中被标记为用药错误率过高的试验比例随预测概率递增而单调上升。研究证明,利用启动前信息可在试验层面预测用药错误风险。概率校准对于将模型输出转化为可靠、可解释的风险类别至关重要;简单多模态融合即可获得性能提升,无需复杂架构。本研究提出一个可复现、可扩展的机器学习框架,支持临床研究中主动、基于风险的质量管理。
原文摘要 · Abstract (English)
Objective: The objective of this study is to develop a machine learning (ML)-based framework for early risk stratification of clinical trials (CTs) according to their likelihood of exhibiting a high rate of dosing errors, using information available prior to trial initiation. Materials and Methods: We constructed a dataset from ClinicalTrials.gov comprising 42,112 CTs. Structured, semi-structured trial data, and unstructured protocol-related free-text data were extracted. CTs were assigned binary labels indicating elevated dosing error rate, derived from adverse event reports, MedDRA terminology, and Wilson confidence intervals. We evaluated an XGBoost model trained on structured features, a ClinicalModernBERT model using textual data, and a simple late-fusion model combining both modalities. Post-hoc probability calibration was applied to enable interpretable, trial-level risk stratification. Results: The late-fusion model achieved the highest AUC-ROC (0.862). Beyond discrimination, calibrated outputs enabled robust stratification of CTs into predefined risk categories. The proportion of trials labeled as having an excessively high dosing error rate increased monotonically across higher predicted risk groups and aligned with the corresponding predicted probability ranges. Discussion: These findings indicate that dosing error risk can be anticipated at the trial level using pre-initiation information. Probability calibration was essential for translating model outputs into reliable and interpretable risk categories, while simple multimodal integration yielded performance gains without requiring complex architectures. Conclusion: This study introduces a reproducible and scalable ML framework for early, trial-level risk stratification of CTs at risk of high dosing error rates, supporting proactive, risk-based quality management in clinical research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。