用报告生成3D CT影像,兼顾一致性与多样性。
CTFlow: Video-Inspired Latent Flow Matching for 3D CT Synthesis
- 基于文本和自回归机制逐步生成CT切片,保持内存可控。
- 在FID、FVD、CLIP等指标上优于现有模型。
- 适合医学数据增强与隐私保护场景使用。
基于临床报告生成完整3D CT体积具有加速研究的潜力,可通过数据增强、隐私保护合成及降低监管约束来释放患者数据价值,同时保留诊断信号。随着大规模配对数据集CT-RATE的发布,训练大型文本条件化CT生成模型成为可能。本文提出CTFlow,一个0.5B参数的潜在流匹配变换器模型,基于临床报告生成3D CT。我们采用FLUX中的A-VAE定义潜在空间,并使用CT-Clip编码器处理临床报告。为生成连贯的完整体积并控制内存开销,采用定制自回归策略:仅凭文本预测初始切片序列,随后结合已生成切片与文本逐步预测后续序列。在对比当前最优生成模型的评估中,我们的方法在时间一致性、图像多样性及文本-图像对齐方面表现更优,各项指标(FID、FVD、IS、CLIP分数)均取得提升。
原文摘要 · Abstract (English)
Generative modelling of entire CT volumes conditioned on clinical reports has the potential to accelerate research through data augmentation, privacy-preserving synthesis and reducing regulator-constraints on patient data while preserving diagnostic signals. With the recent release of CT-RATE, a large-scale collection of 3D CT volumes paired with their respective clinical reports, training large text-conditioned CT volume generation models has become achievable. In this work, we introduce CTFlow, a 0.5B latent flow matching transformer model, conditioned on clinical reports. We leverage the A-VAE from FLUX to define our latent space, and rely on the CT-Clip text encoder to encode the clinical reports. To generate consistent whole CT volumes while keeping the memory constraints tractable, we rely on a custom autoregressive approach, where the model predicts the first sequence of slices of the volume from text-only, and then relies on the previously generated sequence of slices and the text, to predict the following sequence. We evaluate our results against state-of-the-art generative CT model, and demonstrate the superiority of our approach in terms of temporal coherence, image diversity and text-image alignment, with FID, FVD, IS scores and CLIP score.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。