arXiv:2410.00536eess.IVcs.AI2024-10中稿 · MLMI, MICCAI被引 2

用时空变压器自动评估肠镜视频中的溃疡性结肠炎严重程度

Arges: Spatio-Temporal Transformer for Ulcerative Colitis Severity Assessment in Endoscopy Videos

  • 基于带位置编码的Transformer,融合帧间时空信息进行评分
  • 在四项评分指标上显著优于现有方法,最高提升18.8%
  • 适用于临床试验数据,对新病程评分有良好泛化能力

准确评估溃疡性结肠炎(UC)内镜视频中的疾病严重程度对临床试验药物疗效评价至关重要。严重程度通常由梅奥内镜亚评分(MES)和溃疡性结肠炎内镜指数(UCEIS)衡量。然而,专家标注耗时且存在阅片者差异,自动化可缓解此问题。现有基于帧级标签的全监督方法受限于临床试验中普遍存在的视频级标签。基于CNN的弱监督模型(WSL)虽支持端到端训练,但泛化能力差,且忽略对评分至关重要的时空信息。为此,我们提出“Arges”——一种利用带有位置编码的Transformer,从帧特征中整合时空信息以估计内镜视频疾病严重程度的深度学习框架。特征提取来自一个在多中心临床试验大规模数据集(6100万帧,3927个视频)上预训练的基础模型(ArgesFM)。我们在四项UC严重程度评分上进行评估,包括MES及三个UCEIS分项评分。测试集结果表明,相比最先进方法,MES的F1分数提升4.1%,三项UCEIS分项评分分别提升18.8%、6.6%、3.8%。在未见过的临床试验数据上的前瞻性验证进一步证明了模型的良好泛化能力。

原文摘要 · Abstract (English)

Accurate assessment of disease severity from endoscopy videos in ulcerative colitis (UC) is crucial for evaluating drug efficacy in clinical trials. Severity is often measured by the Mayo Endoscopic Subscore (MES) and Ulcerative Colitis Endoscopic Index of Severity (UCEIS) score. However, expert MES/UCEIS annotation is time-consuming and susceptible to inter-rater variability, factors addressable by automation. Automation attempts with frame-level labels face challenges in fully-supervised solutions due to the prevalence of video-level labels in clinical trials. CNN-based weakly-supervised models (WSL) with end-to-end (e2e) training lack generalization to new disease scores and ignore spatio-temporal information crucial for accurate scoring. To address these limitations, we propose "Arges", a deep learning framework that utilizes a transformer with positional encoding to incorporate spatio-temporal information from frame features to estimate disease severity scores in endoscopy video. Extracted features are derived from a foundation model (ArgesFM), pre-trained on a large diverse dataset from multiple clinical trials (61M frames, 3927 videos). We evaluate four UC disease severity scores, including MES and three UCEIS component scores. Test set evaluation indicates significant improvements, with F1 scores increasing by 4.1% for MES and 18.8%, 6.6%, 3.8% for the three UCEIS component scores compared to state-of-the-art methods. Prospective validation on previously unseen clinical trial data further demonstrates the model's successful generalization.

医学影像内镜分析时空建模评分自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。