arXiv:2603.23041cs.CVcs.AI2026-03

分段生成肺部CT,提升图像质量并降低计算成本。

HUydra: Full-Range Lung CT Synthesis via Multiple HU Interval Generative Modelling

  • 按密度区间逐段生成CT,再融合成完整图像。
  • FID提升6.2%,多指标优于传统方法。
  • 适合需要高质量医学影像的科研与临床团队。

当前医学影像中计算机辅助诊断模型的部署与验证面临数据稀缺的瓶颈。针对全球最常见的肺癌,数据不足会延缓诊断并影响患者预后。生成式AI为该问题提供新思路,但全范围亨氏单位(HU)肺部CT分布复杂,建模难度大且计算开销高。本文提出一种新分解策略:不一次性建模整个HU范围,而是分区间逐个生成,再通过可学习重建网络合并为完整扫描。引入多头与多解码器结构以更好捕捉纹理并保持解剖一致性,其中多头VQVAE表现最佳。定量评估显示,该方法显著优于传统2D全范围基线,在所有HU区间上均实现更高的FID(改善6.2%)、MMD、精度与召回率。最优模型在提升视觉保真度和多样性的同时,还降低了模型复杂度与计算成本。本工作建立了一种结构感知的医学图像生成新范式,使生成建模更契合临床解读需求。

原文摘要 · Abstract (English)

Currently, a central challenge and bottleneck in the deployment and validation of computer-aided diagnosis (CAD) models within the field of medical imaging is data scarcity. For lung cancer, one of the most prevalent types worldwide, limited datasets can delay diagnosis and have an impact on patient outcome. Generative AI offers a promising solution for this issue, but dealing with the complex distribution of full Hounsfield Unit (HU) range lung CT scans is challenging and remains as a highly computationally demanding task. This paper introduces a novel decomposition strategy that synthesizes CT images one HU interval at a time, rather than modelling the entire HU domain at once. This framework focuses on training generative architectures on individual tissue-focused HU windows, then merges their output into a full-range scan via a learned reconstruction network that effectively reverses the HU-windowing process. We further propose multi-head and multi-decoder models to better capture textures while preserving anatomical consistency, with a multi-head VQVAE achieving the best performance for the generative task. Quantitative evaluation shows this approach significantly outperforms conventional 2D full-range baselines, achieving a 6.2% improvement in FID and superior MMD, Precision, and Recall across all HU intervals. The best performance is achieved by a multi-head VQVAE variant, demonstrating that it is possible to enhance visual fidelity and variability while also reducing model complexity and computational cost. This work establishes a new paradigm for structure-aware medical image synthesis, aligning generative modelling with clinical interpretation.

CT生成生成模型医学影像VQVAE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。