arXiv:2502.00346physics.med-phcs.AI2025-02被引 4

用强化学习自动规划前列腺癌放疗,仅需一例训练即达顶尖质量。

Actor Critic with Experience Replay-based automatic treatment planning for prostate cancer intensity modulated radiotherapy

  • 基于ACER框架的强化学习模型,通过经验回放优化放疗参数。
  • 单例训练后93%计划获满分9分,平均分达8.93±0.27,显著优于原始方案。
  • 适用于多种数据集,且对对抗攻击具备强鲁棒性,适合临床部署。

背景:调强放射治疗(IMRT)实时规划因复杂射野交互而困难。人工智能虽提升自动化水平,但现有模型依赖大量高质量数据,泛化能力有限。深度强化学习(DRL)通过模拟人类试错过程提供新路径。目的:构建基于随机策略的DRL代理,实现高效训练、广泛适用且对抗攻击下稳健的自动放疗规划,采用快速梯度符号法(FGSM)测试鲁棒性。方法:使用带经验回放的演员-评论家(ACER)架构,代理在逆向规划中调节治疗规划参数(TPPs),输入为剂量体积直方图(DVHs)。模型在单个前列腺癌病例上训练,于两个独立病例验证,并在三个数据集共300+计划上测试。计划质量以ProKnow分数评估,鲁棒性通过对抗攻击测试。结果:尽管仅基于单一病例训练,模型表现出良好泛化能力。ACER规划前平均计划分为6.20±1.84;规划后93.09%案例达满分9分,平均分8.93±0.27。代理有效优先优化最优TPP,且对对抗攻击保持稳健。结论:基于ACER的DRL代理实现了高效、高质量的前列腺癌IMRT自动规划,展现出强泛化性与鲁棒性。

原文摘要 · Abstract (English)

Background: Real-time treatment planning in IMRT is challenging due to complex beam interactions. AI has improved automation, but existing models require large, high-quality datasets and lack universal applicability. Deep reinforcement learning (DRL) offers a promising alternative by mimicking human trial-and-error planning. Purpose: Develop a stochastic policy-based DRL agent for automatic treatment planning with efficient training, broad applicability, and robustness against adversarial attacks using Fast Gradient Sign Method (FGSM). Methods: Using the Actor-Critic with Experience Replay (ACER) architecture, the agent tunes treatment planning parameters (TPPs) in inverse planning. Training is based on prostate cancer IMRT cases, using dose-volume histograms (DVHs) as input. The model is trained on a single patient case, validated on two independent cases, and tested on 300+ plans across three datasets. Plan quality is assessed using ProKnow scores, and robustness is tested against adversarial attacks. Results: Despite training on a single case, the model generalizes well. Before ACER-based planning, the mean plan score was 6.20$\pm$1.84; after, 93.09% of cases achieved a perfect score of 9, with a mean of 8.93$\pm$0.27. The agent effectively prioritizes optimal TPP tuning and remains robust against adversarial attacks. Conclusions: The ACER-based DRL agent enables efficient, high-quality treatment planning in prostate cancer IMRT, demonstrating strong generalizability and robustness.

放疗规划强化学习ACER前列腺癌

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。