arXiv:2511.18595cs.CVcs.AI2025-11被引 2

首个针对胶质母细胞瘤随访MRI的分阶段深度学习基准测试

Timepoint-Specific Benchmarking of Deep Learning Models for Glioblastoma Follow-Up MRI

  • 按时间点独立评估11类模型,统一质量控制流程训练
  • 第二阶段诊断准确率提升,但整体区分能力仍有限
  • Mamba+CNN组合兼顾效率与性能,适合临床部署

区分胶质母细胞瘤治疗后的真性进展(TP)与假性进展(PsP)极具挑战,尤其在早期随访中。本文基于Burdenko GBM Progression队列(n=180)开展首个分阶段、横断面的深度学习模型基准测试。对放疗后不同时间点扫描独立分析,检验模型性能是否随时间变化。采用统一的质量控制驱动流程,训练了11类代表性深度学习模型(CNN、LSTM、混合模型、Transformer及选择性状态空间模型),并使用患者级交叉验证。全阶段准确率相近(约0.70–0.74),但第二阶段的F1和AUC普遍提升,表明后期可分性增强。其中,Mamba+CNN混合模型表现最佳,兼顾准确率与效率;Transformer变体虽有良好AUC,但计算开销大;轻量级CNN效率高但可靠性较差。性能受批量大小影响明显,凸显标准化训练协议的重要性。总体区分能力仍较弱,反映TP与PsP本质区分难度及数据集样本不平衡问题。研究建立阶段感知基准,推动未来工作引入纵向建模、多序列MRI及更大规模多中心队列。

原文摘要 · Abstract (English)

Differentiating true tumor progression (TP) from treatment-related pseudoprogression (PsP) in glioblastoma remains challenging, especially at early follow-up. We present the first stage-specific, cross-sectional benchmarking of deep learning models for follow-up MRI using the Burdenko GBM Progression cohort (n = 180). We analyze different post-RT scans independently to test whether architecture performance depends on time-point. Eleven representative DL families (CNNs, LSTMs, hybrids, transformers, and selective state-space models) were trained under a unified, QC-driven pipeline with patient-level cross-validation. Across both stages, accuracies were comparable (~0.70-0.74), but discrimination improved at the second follow-up, with F1 and AUC increasing for several models, indicating richer separability later in the care pathway. A Mamba+CNN hybrid consistently offered the best accuracy-efficiency trade-off, while transformer variants delivered competitive AUCs at substantially higher computational cost and lightweight CNNs were efficient but less reliable. Performance also showed sensitivity to batch size, underscoring the need for standardized training protocols. Notably, absolute discrimination remained modest overall, reflecting the intrinsic difficulty of TP vs. PsP and the dataset's size imbalance. These results establish a stage-aware benchmark and motivate future work incorporating longitudinal modeling, multi-sequence MRI, and larger multi-center cohorts.

医学影像深度学习胶质瘤时间建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。