用深度强化学习自动重规划头颈癌质子治疗,提升疗效并省去人工耗时。
Patient-Specific Deep Reinforcement Learning for Automatic Replanning in Head-and-Neck Cancer Proton Therapy
- 为每位患者定制强化学习模型,利用模拟解剖变化数据训练智能体。
- 重规划后计划评分从120.78升至141.50,优于人工重规划的136.32。
- 适合需要高效自适应质子治疗的临床团队和研究者参考。
头颈癌质子调强治疗(IMPT)过程中解剖结构变化会导致布拉格峰偏移,可能造成肿瘤剂量不足和危及器官过量照射。当前手动重规划耗时且资源密集。本文提出一种针对患者的深度强化学习(DRL)框架,通过基于150点计划质量评分的奖励设计,解决临床目标间的权衡问题。将计划过程建模为强化学习任务,智能体学习调整优化优先级以最大化计划质量。不同于群体方法,该框架使用每位患者的计划CT及模拟解剖变化(肿瘤进展与退缩)数据进行个体化训练,利用治疗周期中的解剖相似性实现有效适应。采用Deep Q-Network与Proximal Policy Optimization两种算法,以剂量体积直方图(DVHs)为状态表示,22维动作空间用于优先级调整。在8名头颈癌患者的真实重规划CT数据上评估,两种智能体将初始计划评分从120.78±17.18提升至139.59±5.50(DQN)和141.50±4.69(PPO),显著优于人工重规划的136.32±4.79。临床验证表明,改进效果转化为更优的肿瘤覆盖与危及器官保护,应对几何与剂量复杂性具有潜力,为离线自适应及在线自适应质子治疗提供高效解决方案。
原文摘要 · Abstract (English)
Anatomical changes during intensity-modulated proton therapy (IMPT) for head-and-neck cancer (HNC) can shift Bragg peaks, risking tumor underdosing and organ-at-risk overdosing. Treatment replanning is often required to maintain clinically acceptable treatment quality. However, current manual replanning processes are resource-intensive and time-consuming. We propose a patient-specific deep reinforcement learning (DRL) framework for automated IMPT replanning, with a reward-shaping mechanism based on a $150$-point plan quality score addressing competing clinical objectives. We formulate the planning process as a reinforcement learning problem where agents learn control policies to adjust optimization priorities, maximizing plan quality. Unlike population-based approaches, our framework trains agents for each patient using their planning Computed Tomography (CT) and augmented anatomies simulating anatomical changes (tumor progression and regression). This patient-specific approach leverages anatomical similarities along the treatment course, enabling effective plan adaptation. We implemented two DRL algorithms, Deep Q-Network and Proximal Policy Optimization, using dose-volume histograms (DVHs) as state representations and a $22$-dimensional action space of priority adjustments. Evaluation on eight HNC patients using actual replanning CT data showed that both agents improved initial plan scores from $120.78 \pm 17.18$ to $139.59 \pm 5.50$ (DQN) and $141.50 \pm 4.69$ (PPO), surpassing the replans manually generated by a human planner ($136.32 \pm 4.79$). Clinical validation confirms that improvements translate to better tumor coverage and OAR sparing across diverse anatomical changes. This work highlights DRL's potential in addressing geometric and dosimetric complexities of adaptive proton therapy, offering efficient offline adaptation solutions and advancing online adaptive proton therapy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。