用强化学习提升3D肺部CT生成的临床可控性,真实度和控制力双突破。
CONFLUX: A Latent Diffusion Model for 3D Chest-CT Synthesis with RL Post-Training

- 基于潜空间扩散模型,结合自适应归一化实现多属性条件生成
- 生成体积在三平面FID上达32.3,优于基线74.6,且临床属性控制更精准
- 强化学习后处理使生成结果与真实扫描的诊断可靠性差距缩小47%
可控的3D医学图像生成模型可合成具有指定临床特征的体数据,但需同时满足高保真、原生3D和条件忠实。本文提出CONFLUX,一种用于胸部计算机断层扫描(CT)的潜扩散模型:3D变分自编码器压缩体数据,修正流变压器在潜空间生成。生成过程通过结构化放射学元数据(18种异常发现、性别、年龄、重建核)进行条件控制,采用自适应层归一化。该模型在三平面弗雷谢距离(FID 32.3)上显著优于现有基线(74.6),并实现对临床属性的直接控制。为进一步增强控制能力,引入在线强化学习后训练阶段(组相对策略优化),以分类器从生成体中准确恢复请求异常的可靠性为奖励。经独立分类器评估,后训练使生成结果与真实扫描的诊断可靠性差距减少47%。我们发布了模型及约20万张带有条件元数据的合成胸部CT数据集,覆盖广泛临床发现。
原文摘要 · Abstract (English)
Controllable generative models of 3D medical images can synthesize volumes with specified clinical attributes, but this demands samples that are simultaneously high-fidelity, natively 3D, and faithful to the requested conditioning. We present CONFLUX, a latent diffusion model for chest computed tomography (CT): a 3D variational autoencoder compresses each volume, and a rectified-flow transformer generates in the latent space. Generation is conditioned on structured radiological metadata (18 abnormality findings, sex, age, and reconstruction kernel) through adaptive layer normalization. The model leads strong volumetric baselines on tri-planar Frechet distance (FID 32.3 vs. 74.6 for MAISI) while exposing direct control over clinical attributes. To strengthen that control we add an online reinforcement-learning post-training stage (group-relative policy optimization) that rewards how reliably a classifier recovers the requested findings from each generated volume. Judged by a separate, independent classifier, post-training removes 47% of the shortfall relative to real-scan reliability. We release the model and a ~200k synthetic chest-CT dataset with conditioning metadata spanning a wide variety of clinical findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。