对比7种生成模型在3D医学影像跨模态转换中的表现,发现GANs更优且合成图像难辨真伪。
Cross Modality Image Translation In Medical Imaging Using Generative Frameworks

- 构建标准化3D医学影像转换框架,统一训练、推理与评估流程
- SRGAN在11个数据集77次实验中显著优于其他模型,尤其在结构保持上
- 合成图像临床可接受度高,但小病灶和定量强度仍存缺陷,适合研究者参考
医学图像到图像(I2I)转换可实现虚拟扫描,即无需额外采集即可从源模态合成目标模态图像。尽管关注度上升,多数方法局限于2D切片,评估任务孤立且实验设置各异,缺乏临床验证。本文首次提出可复现的标准化3D I2I转换评估框架,涵盖预处理、划分、推理与多层级评估,覆盖头颈、肺、盆腔三个解剖区域,四种转换方向(CBCT→CT、MRI→CT、CT→PET、T2→T2-FLAIR),共11个数据集,77次实验。比较七种生成模型:三种GAN(Pix2Pix、CycleGAN、SRGAN)和四种潜在生成模型(LDM、LDM+ControlNet、Brownian Bridge、Flow Matching)。结果显示,所有任务中GAN均优于潜在生成模型,其中SRGAN表现显著更优。病灶级分析表明,所有模型对小病灶处理不佳;在CT→PET转换中,模型更可靠地保留病灶形状而非摄取强度。17名医生(含15名放射科医师)参与视觉图灵测试,分类准确率仅56.7%(接近随机水平),证实合成图像与真实图像难以区分,但量化指标与临床偏好存在分离现象。
原文摘要 · Abstract (English)
Medical image-to-image (I2I) translation enables virtual scanning, i.e. the synthesis of a target imaging modality from a source one without additional acquisitions. Despite growing interest, most proposed methods operate on 2D slices, are evaluated on isolated tasks with different experimental set-ups and lack clinical validation. The primary contribution of this work is a reproducible, standardized comparative evaluation of 3D I2I translation methods in oncological imaging, designed to standardize preprocessing, splitting, inference, and multi-level evaluation across heterogeneous clinical tasks. Within this framework, we compare seven generative models, three Generative Adversarial Networks (GANs: Pix2Pix, CycleGAN, SRGAN) and four latent generative models (Latent Diffusion Model, Latent Diffusion Model+ControlNet, Brownian Bridge, Flow Matching), across eleven datasets spanning three anatomical regions (head/neck, lung, pelvis) and four translation directions (cone-beam CT to CT, MRI to CT, CT to PET, MRI T2-weighted to T2-FLAIR), for a total of 77 experiments under uniform training, inference, and evaluation conditions. The results show that GANs outperform latent generative models across all tasks, with SRGAN achieving statistically significant superiority. Our lesion-level analysis reveals that all models struggle with small lesions and that, in CT to PET synthesis, models reproduce lesion shape more reliably than absolute uptake-related intensity. We also performed a Visual Turing test administered to 17 physicians, including 15 radiologists, which shows near-chance classification accuracy (56.7%), confirming that synthetic volumes are largely indistinguishable from real acquisitions, while exposing a dissociation between quantitative metrics and clinical preference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。