arXiv:2506.11957physics.med-phcs.LG2025-06被引 2

用强化学习自动设计宫颈癌放疗方案,效率与一致性显著提升。

Automated Treatment Planning for Interstitial HDR Brachytherapy for Locally Advanced Cervical Cancer using Deep Reinforcement Learning

  • 分两阶段:先用深度Q网络选参数,再用自定义优化器算照射时间
  • 新方法在未见患者上得分93.89%,优于临床计划的91.86%
  • 保持靶区全覆盖,多数情况减少热点,适合放疗自动化研究者

高剂量率(HDR)腔内近距离放疗在局部晚期宫颈癌治疗中至关重要,但高度依赖人工规划。本研究旨在开发一种整合强化学习(RL)与剂量优化的全自动HDR近距离放疗规划框架,以提高计划的一致性与效率。提出分层两阶段自动规划框架:第一阶段采用基于深度Q网络(DQN)的强化学习代理,迭代选择控制靶区覆盖与危及器官(OAR)保护权衡的治疗规划参数(TPPs),状态表示包含剂量体积直方图(DVH)指标与当前TPP值,奖励函数融合临床剂量目标与安全约束,涵盖靶区的D90、V150、V200,以及所有相关危及器官(膀胱、直肠、乙状结肠、小肠、大肠)的D2cc;第二阶段采用定制的基于Adam的优化器,利用临床导向损失函数计算所选TPPs对应的驻留时间分布。该框架在具有复杂插植几何结构的患者队列上进行评估。结果表明,该框架成功在不同患者解剖结构下学习到临床有意义的TPP调整。对于未见过的测试患者,强化学习自动化规划方法平均得分为93.89%,优于临床计划的91.86%。值得注意的是,分数提升的同时,均保持了完整的靶区覆盖,并在多数情况下减少了靶区热点。

原文摘要 · Abstract (English)

High-dose-rate (HDR) brachytherapy plays a critical role in the treatment of locally advanced cervical cancer but remains highly dependent on manual treatment planning expertise. The objective of this study is to develop a fully automated HDR brachytherapy planning framework that integrates reinforcement learning (RL) and dose-based optimization to generate clinically acceptable treatment plans with improved consistency and efficiency. We propose a hierarchical two-stage autoplanning framework. In the first stage, a deep Q-network (DQN)-based RL agent iteratively selects treatment planning parameters (TPPs), which control the trade-offs between target coverage and organ-at-risk (OAR) sparing. The agent's state representation includes both dose-volume histogram (DVH) metrics and current TPP values, while its reward function incorporates clinical dose objectives and safety constraints, including D90, V150, V200 for targets, and D2cc for all relevant OARs (bladder, rectum, sigmoid, small bowel, and large bowel). In the second stage, a customized Adam-based optimizer computes the corresponding dwell time distribution for the selected TPPs using a clinically informed loss function. The framework was evaluated on a cohort of patients with complex applicator geometries. The proposed framework successfully learned clinically meaningful TPP adjustments across diverse patient anatomies. For the unseen test patients, the RL-based automated planning method achieved an average score of 93.89%, outperforming the clinical plans which averaged 91.86%. These findings are notable given that score improvements were achieved while maintaining full target coverage and reducing CTV hot spots in most cases.

放疗规划强化学习自动化宫颈癌

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。