用AI自动优化质子治疗计划,速度快且效果更好。
A learning-driven automatic planning framework for proton PBS treatments of H&N cancers
- 用学习型逆优化算法自动调整治疗目标参数
- 平均2.55小时生成高质量计划,比人工快36%以上
- 适合需要快速精准治疗的头颈部癌患者
针对头颈部癌症的质子笔形束扫描(PBS)治疗计划涉及多个相互冲突的目标,需反复调整目标参数以平衡临床需求。本文提出一种基于学习的逆优化器,并集成到近端策略优化(PPO)框架中,实现对不同治疗需求患者的全自动高质量计划生成。逆优化器采用学习即优化(L2O)方法,通过学习特定任务数据分布预测更新步长。首次将大语言模型中的长上下文处理技术应用于L2O,突破了现有方法在大规模变量同时优化上的可扩展性限制。PPO框架作为外层虚拟规划器,通过策略网络自动调节目标参数;内层L2O逆优化器根据优化后的目标计算可执行的点剂量监视单元(MU)值。此外,训练了一个Swin UnetR剂量预测器,结合处方和束流特异性信息估计初始目标参数。实验共使用97例双侧或同侧头颈部癌症患者数据进行训练与测试。相比二阶梯度方法,该L2O优化器在逆优化的有效性和效率上分别提升22.97%和36.41%;结合PPO框架后,平均计划生成时间仅2.55小时,满足临床要求,且靶区覆盖优于、危及器官保护不低于人工计划。
原文摘要 · Abstract (English)
Proton pencil beam scanning (PBS) treatment planning for head & neck (H&N) cancers involves numerous conflicting objectives, requiring iterative objective parameter adjustments to balance multiple clinical goals. We propose a learning-driven inverse optimizer and integrate it into a proximal policy optimization (PPO)-based planning framework to automatically generate high-quality plans for patients with diverse treatment requirements. The inverse optimizer is a learning-to-optimize (L2O) method that predicts update steps by learning from task-specific data distributions. For the first time, long-context processing techniques developed for large language models (LLMs) are utilized to address the scalability limitations of existing L2O methods, enabling simultaneous optimization over a substantially large set of variables. The PPO framework functions as an outer-loop virtual planner, autonomously adjusting objective parameters through a policy network, and the inner-loop L2O inverse optimizer computes machine-deliverable spot monitor unit (MU) values based on the PPO-refined objectives. Moreover, a Swin UnetR dose predictor is trained with prescription- and beam-specific information to estimate the initial objective parameters. In our experiments, total 97 patients with bilateral or ipsilateral H&N cancers are collected for training and testing. Compared with the second-order gradient-based methods, our L2O optimizer improves the effectiveness and efficiency of the time-consuming inverse optimization by 22.97% and 36.41%, respectively, and in conjunction with the PPO-based virtual planner, plans are generated within clinically acceptable times, i.e. 2.55 hours in average, and shows improved or comparable organs-at-risk sparing with superior target coverage compared with human-generated plans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。