arXiv:2606.00922physics.med-phcs.RO2026-06

用强化学习指导大模型,实现无需人工干预的放疗计划自动化。

A Machine-to-Machine Knowledge-Guided LLM Agent for Generalizable Radiotherapy Treatment Planning

论文配图:A Machine-to-Machine Knowledge-Guided LLM Agent for Generalizable Radiotherapy Treatment Planning
图 1 · 摘自论文原文
  • 通过强化学习提取参数知识,注入大模型实现自主迭代规划。
  • 相比无引导方案,迭代次数显著减少,且在三种场景下均达最优计划评分。
  • 能适应不同解剖结构和初始条件,适合临床真实环境部署。

本文提出一种机器间(M2M)知识引导的大语言模型(LLM)框架,用于自动化放疗治疗计划。该框架将深度强化学习(DRL)代理发现的治疗计划参数(TPP)分布知识,通过上下文学习传递给LLM代理,实现无需人工干预的自主迭代规划。传统基于LLM的规划常缺乏物理直觉且难以收敛,而引入DRL引导可约束代理在物理可行参数空间内搜索。实验在三种不同场景中评估:基础前列腺病例、具有更高危器官(OAR)约束的复杂前列腺病例,以及肝脏病例。结果表明,引导后的LLM代理始终达到最优计划评分,同时大幅减少迭代次数。最终的TPP配置分析显示,代理成功学习到目标的层次化优先级,重建了参数调整与剂量学结果间的逻辑“因果”关系。关键的是,该原型框架展现出强泛化能力,在不同患者解剖、治疗部位及初始计划质量下均保持高计划质量。通过融合DRL的专项优化与LLM的自适应推理,该M2M框架为通用自主治疗计划建立了可扩展基础,有望在真实临床环境中提升实践效率。

原文摘要 · Abstract (English)

In this work, we propose a prototype machine-to-machine (M2M) knowledge-guided Large Language Model (LLM) framework for automated radiotherapy treatment planning. In the proposed paradigm, Treatment Planning Parameter (TPP) distribution knowledge discovered by a Deep Reinforcement Learning (DRL) agent is transferred to an LLM agent through in-context learning, enabling autonomous iterative planning without human intervention. While standard LLM-based planning often lacks physical intuition and struggles with convergence, the integration of DRL-derived guidance constrains the agent to a physically valid parameter space. Experimental evaluations are performed across three diverse planning scenarios: basic prostate cases, complex prostate configurations with increased organ-at-risk (OAR) constraints, and liver cases. The evaluation results demonstrate that the guided LLM agent consistently achieves optimal planning scores while significantly reducing the number of iterations compared to unguided planning. Analysis of the final TPP configurations reveals that the agent successfully learns a hierarchical priority of objectives, effectively restoring a logical "cause-and-effect" relationship between parameter tuning and dosimetric outcomes. Crucially, this prototype framework exhibits robust generalizability, maintaining high planning quality regardless of specific patient anatomy, treatment site, or initial plan quality. By bridging the specialized optimization of DRL with the adaptive reasoning of LLMs, this M2M framework establishes a scalable foundation towards generalizable autonomous treatment planning, ultimately benefiting clinical practice in realistic environments.

放疗计划大模型强化学习自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。