arXiv:2506.09805physics.med-phcs.LG2025-06被引 1

用强化学习自动规划前列腺癌放疗,更精准且少用穿刺针。

Automatic Treatment Planning using Reinforcement Learning for High-dose-rate Prostate Brachytherapy

  • 通过强化学习逐针优化针位与照射时间
  • 计划质量相当,热点更少,平均少用2根针
  • 适合想标准化放疗规划的临床团队

目的:高剂量率(HDR)前列腺腔内放疗中,穿刺针布局依赖医生经验。本文研究利用强化学习(RL)在术前规划阶段根据患者解剖结构自动生成针位与照射时间,以缩短手术时间并保证计划质量一致性。方法:训练一个强化学习代理,观察环境后调整选定针的位置和所有照射时间,以最大化预设奖励函数;调整完毕后,再处理下一针,直至所有针完成调整。重复多轮直至达到最大轮数。本研究纳入11例前列腺HDR加强治疗患者数据(1例用于训练,10例用于测试)。将强化学习生成的计划与临床实际计划(真实基准)的剂量学指标和使用针数进行比较。结果:强化学习计划与临床计划在前列腺覆盖度(前列腺V100)和直肠受量(直肠D2cc)方面无统计学差异,但强化学习计划的前列腺热点(前列腺V150)和尿道受量(尿道D20%)显著更低。此外,强化学习计划平均比临床计划少使用2根针。结论:这是首个证明强化学习可自主生成临床可行的HDR前列腺腔内放疗计划的研究。该方法在保持或提升计划质量的同时减少针数,具备低数据需求和强泛化能力,有潜力实现放疗规划标准化,降低临床差异,改善患者预后。

原文摘要 · Abstract (English)

Purpose: In high-dose-rate (HDR) prostate brachytherapy procedures, the pattern of needle placement solely relies on physician experience. We investigated the feasibility of using reinforcement learning (RL) to provide needle positions and dwell times based on patient anatomy during pre-planning stage. This approach would reduce procedure time and ensure consistent plan quality. Materials and Methods: We train a RL agent to adjust the position of one selected needle and all the dwell times on it to maximize a pre-defined reward function after observing the environment. After adjusting, the RL agent then moves on to the next needle, until all needles are adjusted. Multiple rounds are played by the agent until the maximum number of rounds is reached. Plan data from 11 prostate HDR boost patients (1 for training, and 10 for testing) treated in our clinic were included in this study. The dosimetric metrics and the number of used needles of RL plan were compared to those of the clinical results (ground truth). Results: On average, RL plans and clinical plans have very similar prostate coverage (Prostate V100) and Rectum D2cc (no statistical significance), while RL plans have less prostate hotspot (Prostate V150) and Urethra D20% plans with statistical significance. Moreover, RL plans use 2 less needles than clinical plan on average. Conclusion: We present the first study demonstrating the feasibility of using reinforcement learning to autonomously generate clinically practical HDR prostate brachytherapy plans. This RL-based method achieved equal or improved plan quality compared to conventional clinical approaches while requiring fewer needles. With minimal data requirements and strong generalizability, this approach has substantial potential to standardize brachytherapy planning, reduce clinical variability, and enhance patient outcomes.

放疗规划强化学习前列腺癌自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。