AI辅助医生读片可稳定提升临床试验可靠性,即使模型表现差也能保持结果可信。
The Framework That Survives Bad Models: Human-AI Collaboration For Clinical Trials
- AI作为辅助读片工具,不替代医生决策
- 在坏模型干扰下仍能准确评估疾病进展
- 适合需要高可靠性的医学影像临床试验
人工智能在临床试验中具有巨大潜力,可用于患者招募、终点评估和治疗反应预测。然而,缺乏保障的AI部署存在显著风险,尤其在影响试验结论的患者终点评估中。我们对比了两种AI框架与纯人工评估在医学影像疾病评价中的表现,衡量成本、准确率、鲁棒性和泛化能力。通过注入从随机猜测到简单预测的劣质模型,对框架进行压力测试,确保即使在严重模型退化情况下,治疗效果估计依然有效。研究基于两项随机对照试验,终点来自脊柱X光片。结果表明,将AI作为辅助读片者(AI-SR)是最优方案,能在不同模型类型下满足所有评估标准,持续提供可靠的疾病评估,保持临床试验治疗效应估计与结论不变,并在不同人群间保持优势。
原文摘要 · Abstract (English)
Artificial intelligence (AI) holds great promise for supporting clinical trials, from patient recruitment and endpoint assessment to treatment response prediction. However, deploying AI without safeguards poses significant risks, particularly when evaluating patient endpoints that directly impact trial conclusions. We compared two AI frameworks against human-only assessment for medical image-based disease evaluation, measuring cost, accuracy, robustness, and generalization ability. To stress-test these frameworks, we injected bad models, ranging from random guesses to naive predictions, to ensure that observed treatment effects remain valid even under severe model degradation. We evaluated the frameworks using two randomized controlled trials with endpoints derived from spinal X-ray images. Our findings indicate that using AI as a supporting reader (AI-SR) is the most suitable approach for clinical trials, as it meets all criteria across various model types, even with bad models. This method consistently provides reliable disease estimation, preserves clinical trial treatment effect estimates and conclusions, and retains these advantages when applied to different populations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。