arXiv:2507.07811cs.CV2025-07

用视觉变压器预测肺肿瘤运动,患者特异性更准,多患者模型更稳。

Patient-specific vs Multi-Patient Vision Transformer for Markerless Tumor Motion Forecasting

  • 用视觉变压器建模肿瘤运动轨迹,分患者特异与多患者训练两种策略。
  • 患者特异性模型在规划数据上误差低至1.2mm(ADE),多患者模型处理治疗变体更鲁棒。
  • 适合临床快速部署,尤其在影像数据有限时展现实用价值。

准确预测肺肿瘤运动对质子治疗精准给药至关重要。现有无标记方法多依赖深度学习,但基于变压器的架构在此领域尚未探索,尽管其在轨迹预测中表现优异。本研究首次引入视觉变压器(ViT)进行无标记肿瘤运动预测,评估两种训练策略:患者特异性(PS)模型仅使用目标患者规划数据,多患者(MP)模型则在31名患者的计划4DCT生成的数字重建投影(DRRs)上训练,第32名患者用于评估。两者均以每输入16张DRR预测1秒内运动。性能通过平均位移误差(ADE)和最终位移误差(FDE)评估,涵盖计划期(T1)与治疗期(T2)数据。结果表明:在T1数据上,PS模型在大样本(最多25,000张DRR)下优于MP模型(p < 0.05);但在T2数据上,未重新训练的MP模型表现相当,且对分次间解剖变化更具鲁棒性。结论:这是首个将ViT应用于无标记肿瘤运动预测的研究,虽患者特异性模型精度更高,但多患者模型无需再训练即可稳定运行,更适合时间受限的临床场景。

原文摘要 · Abstract (English)

Background: Accurate forecasting of lung tumor motion is essential for precise dose delivery in proton therapy. While current markerless methods mostly rely on deep learning, transformer-based architectures remain unexplored in this domain, despite their proven performance in trajectory forecasting. Purpose: This work introduces a markerless forecasting approach for lung tumor motion using Vision Transformers (ViT). Two training strategies are evaluated under clinically realistic constraints: a patient-specific (PS) approach that learns individualized motion patterns, and a multi-patient (MP) model designed for generalization. The comparison explicitly accounts for the limited number of images that can be generated between planning and treatment sessions. Methods: Digitally reconstructed radiographs (DRRs) derived from planning 4DCT scans of 31 patients were used to train the MP model; a 32nd patient was held out for evaluation. PS models were trained using only the target patient's planning data. Both models used 16 DRRs per input and predicted tumor motion over a 1-second horizon. Performance was assessed using Average Displacement Error (ADE) and Final Displacement Error (FDE), on both planning (T1) and treatment (T2) data. Results: On T1 data, PS models outperformed MP models across all training set sizes, especially with larger datasets (up to 25,000 DRRs, p < 0.05). However, MP models demonstrated stronger robustness to inter-fractional anatomical variability and achieved comparable performance on T2 data without retraining. Conclusions: This is the first study to apply ViT architectures to markerless tumor motion forecasting. While PS models achieve higher precision, MP models offer robust out-of-the-box performance, well-suited for time-constrained clinical settings.

肿瘤运动预测视觉变压器质子治疗医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。