用预训练扩散模型提升放疗剂量预测,跨场景泛化更强。
Any2Any 3D Diffusion Models with Knowledge Transfer: A Radiotherapy Planning Study

- 通过模态感知嵌入实现多模态灵活条件生成,无注意力开销。
- 剂量预测平均绝对误差降至1.93,优于竞赛优胜方案(2.07)。
- 结合临床偏好强化学习,生成结果更贴合机构治疗习惯。
体素级剂量预测是实际放疗规划中的关键挑战,因从头训练的专属模型难以在不同临床环境中泛化。与此同时,视觉领域基于十亿级数据训练的生成模型已取得显著成效。本文提出DiffKT3D,一种统一的Any2Any 3D扩散框架,利用预训练视频扩散模型的先验知识,实现高效且临床上有意义的剂量预测。为支持跨多种临床模态(如CT、解剖结构、身体形态、束流设置等)的灵活条件输入,我们引入无需交叉注意力开销的模态特定嵌入式Any2Any条件范式。此外,设计了一种由临床导向评分卡引导的新型强化学习后训练机制,明确匹配机构治疗偏好。相较于GDP-HMM挑战赛优胜方案,DiffKT3D将体素级平均绝对误差从2.07降至1.93,同时在图像质量和治疗偏好匹配度上表现更优。结果表明,通过模态感知条件与临床对齐的强化学习后训练迁移扩散先验,可为各类临床场景提供鲁棒且泛化的放疗规划解决方案。
原文摘要 · Abstract (English)
Voxel-wise dose prediction is a critical yet challenging task in practical radiotherapy (RT) planning, as bespoke models trained from scratch often struggle to generalize across diverse clinical settings. Meanwhile, generative models trained on billion-scale datasets from vision domains have achieved impressive performance. Herein, we propose DiffKT3D, a unified Any2Any 3D diffusion framework that leverages prior knowledge from pretrained video diffusion models for efficient and clinically meaningful dose prediction. To enable flexible conditioning across multiple clinical modalities (CT, anatomical structures, body, beam settings, etc.), we introduce an Any2Any conditional paradigm utilizing modality-specific embeddings without cross-attention overhead. Further, we design a novel reinforcement learning (RL) post-training mechanism guided by a clinically-informed Scorecard explicitly tailored to institutional treatment preferences. Compared with winner of GDP-HMM challenge, DiffKT3D sets a new state-of-the-art in dose prediction by reducing voxel-level MAE from 2.07 to 1.93. In addition, DiffKT3D achieves superior image quality and preference match. These results demonstrate that transferring diffusion priors via modality-aware conditioning and clinically aligned RL post-training can provide a robust and generalizable solution for RT planning across various clinical scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。