arXiv:2511.02314cs.LGphysics.med-ph2025-11

用多智能体强化学习自动优化碳离子放疗参数,显著提升关键器官保护效果。

Large-scale automatic carbon ion treatment planning for head and neck cancers via parallel multi-agent reinforcement learning

  • 设计并行多智能体强化学习框架,同时调节45个治疗参数
  • 生成计划质量与专家水平相当,5个关键器官受照剂量显著降低
  • 直接对接放疗系统,适合临床自动化放疗规划场景

头颈部癌放疗规划因多个关键器官靠近复杂靶区而困难。调强碳离子治疗(IMCT)虽具优异剂量适形性和器官保护能力,但受限于相对生物效应(RBE)建模,参数调节耗时且依赖经验,常导致次优结果。现有深度学习方法受数据偏差和计划可行性限制,强化学习则难以高效探索高维参数空间。本文提出可扩展的多智能体强化学习(MARL)框架,实现IMCT中45个治疗参数的并行优化。采用中心化训练、分散执行(CTDE)的QMIX结构,结合双DQN、 Dueling DQN与循环编码(DRQN),在高维非平稳环境中实现稳定学习。为提升效率,采用紧凑的历史DVH向量作为状态输入,通过线性动作-价值映射将离散动作转换为均匀参数调整,并设计基于临床意义的分段奖励函数。构建同步多进程工作系统,与PHOENIX TPS对接,实现并行优化与加速数据采集。在10例训练、10例测试的头颈部数据集上,该方法同时调节45个参数,生成计划质量与专家手动方案相当或更优(相对计划得分:RL 85.93±7.85% vs 手动 85.02±6.92%),在五个关键器官上改善显著(p<0.05)。该框架高效探索高维参数空间,通过直接与TSP交互生成临床可接受的IMCT计划,显著提升器官保护效果。

原文摘要 · Abstract (English)

Head-and-neck cancer (HNC) planning is difficult because multiple critical organs-at-risk (OARs) are close to complex targets. Intensity-modulated carbon-ion therapy (IMCT) offers superior dose conformity and OAR sparing but remains slow due to relative biological effectiveness (RBE) modeling, leading to laborious, experience-based, and often suboptimal tuning of many treatment-planning parameters (TPPs). Recent deep learning (DL) methods are limited by data bias and plan feasibility, while reinforcement learning (RL) struggles to efficiently explore the exponentially large TPP search space. We propose a scalable multi-agent RL (MARL) framework for parallel tuning of 45 TPPs in IMCT. It uses a centralized-training decentralized-execution (CTDE) QMIX backbone with Double DQN, Dueling DQN, and recurrent encoding (DRQN) for stable learning in a high-dimensional, non-stationary environment. To enhance efficiency, we (1) use compact historical DVH vectors as state inputs, (2) apply a linear action-to-value transform mapping small discrete actions to uniform parameter adjustments, and (3) design an absolute, clinically informed piecewise reward aligned with plan scores. A synchronous multi-process worker system interfaces with the PHOENIX TPS for parallel optimization and accelerated data collection. On a head-and-neck dataset (10 training, 10 testing), the method tuned 45 parameters simultaneously and produced plans comparable to or better than expert manual ones (relative plan score: RL $85.93\pm7.85%$ vs Manual $85.02\pm6.92%$), with significant (p-value $<$ 0.05) improvements for five OARs. The framework efficiently explores high-dimensional TPP spaces and generates clinically competitive IMCT plans through direct TPS interaction, notably improving OAR sparing.

放疗规划强化学习碳离子治疗多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。