arXiv:2603.06173cs.CV2026-03被引 1

用多尺度奖励学习优化3D医学图像生成,提升细节与结构质量

Optimizing 3D Diffusion Models for Medical Imaging via Multi-Scale Reward Learning

  • 通过强化学习融合2D切片与3D体数据反馈,指导生成
  • 在BraTS和OASIS-1上FID显著降低,下游分类任务性能提升
  • 适合需要高质量医学合成数据的研究者使用

扩散模型已成为3D医学图像生成的强大工具,但标准训练目标与临床相关性之间仍存在差距。本文提出一种基于强化学习(RL)的多尺度反馈方法,以优化3D扩散模型。首先在MRI体积数据上预训练3D扩散模型,建立稳健的生成先验;随后采用近端策略优化(PPO)进行微调,利用结合2D切片评估与3D体分析的新型奖励系统引导生成。该机制使模型同时优化局部纹理细节与全局结构一致性。我们在BraTS 2019和OASIS-1数据集上验证了该框架。结果表明,引入强化学习反馈能有效引导生成过程趋向更高质量分布。定量分析显示,弗雷歇启动距离(FID)显著降低,且合成数据在下游肿瘤与疾病分类任务中表现优于非优化基线。

原文摘要 · Abstract (English)

Diffusion models have emerged as powerful tools for 3D medical image generation, yet bridging the gap between standard training objectives and clinical relevance remains a challenge. This paper presents a method to enhance 3D diffusion models using Reinforcement Learning (RL) with multi-scale feedback. We first pretrain a 3D diffusion model on MRI volumes to establish a robust generative prior. Subsequently, we fine-tune the model using Proximal Policy Optimization (PPO), guided by a novel reward system that integrates both 2D slice-wise assessments and 3D volumetric analysis. This combination allows the model to simultaneously optimize for local texture details and global structural coherence. We validate our framework on the BraTS 2019 and OASIS-1 datasets. Our results indicate that incorporating RL feedback effectively steers the generation process toward higher quality distributions. Quantitative analysis reveals significant improvements in Fréchet Inception Distance (FID) and, crucially, the synthetic data demonstrates enhanced utility in downstream tumor and disease classification tasks compared to non-optimized baselines.

3D生成医学图像强化学习扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。