用强化学习自动规划头颈癌质子治疗,效果媲美专家
Automating proton PBS treatment planning for head and neck cancers using policy gradient-based deep reinforcement learning
- 用PPO算法在连续空间调整治疗目标参数
- 对重要器官保护更好,靶区覆盖不降低
- 首个达到医生水平的质子治疗自动规划模型
头颈癌质子笔形束扫描(PBS)治疗计划制定耗时且依赖经验,涉及大量规划目标。现有深度强化学习方法基于Q-learning和线性加权临床指标,可扩展性差,仅能离散调整有限目标。本文提出基于近端策略优化(PPO)算法和剂量分布奖励函数的自动治疗计划模型。通过经验规则从靶区和危及器官生成辅助规划结构及其目标,输入自研优化引擎生成光斑监测单位(MU)值。训练的策略网络使用PPO在连续动作空间迭代调整目标参数,利用新型剂量分布奖励函数优化质子PBS计划。与人工计划相比,模型生成的计划在同等或更优靶区覆盖下显著改善危及器官保护。肝脏癌额外实验表明该方法可成功泛化至其他部位。据我们所知,这是首个实现头颈癌质子治疗自动规划且达到人类水平性能的DRL模型。
原文摘要 · Abstract (English)
Proton pencil beam scanning (PBS) treatment planning for head and neck (H&N) cancers is a time-consuming and experience-demanding task where a large number of planning objectives are involved. Deep reinforcement learning (DRL) has recently been introduced to the planning processes of intensity-modulated radiation therapy and brachytherapy for prostate, lung, and cervical cancers. However, existing approaches are built upon the Q-learning framework and weighted linear combinations of clinical metrics, suffering from poor scalability and flexibility and only capable of adjusting a limited number of planning objectives in discrete action spaces. We propose an automatic treatment planning model using the proximal policy optimization (PPO) algorithm and a dose distribution-based reward function for proton PBS treatment planning of H&N cancers. Specifically, a set of empirical rules is used to create auxiliary planning structures from target volumes and organs-at-risk (OARs), along with their associated planning objectives. These planning objectives are fed into an in-house optimization engine to generate the spot monitor unit (MU) values. A decision-making policy network trained using PPO is developed to iteratively adjust the involved planning objective parameters in a continuous action space and refine the PBS treatment plans using a novel dose distribution-based reward function. Proton H&N treatment plans generated by the model show improved OAR sparing with equal or superior target coverage when compared with human-generated plans. Moreover, additional experiments on liver cancer demonstrate that the proposed method can be successfully generalized to other treatment sites. To the best of our knowledge, this is the first DRL-based automatic treatment planning model capable of achieving human-level performance for H&N cancers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。