arXiv:2509.14780cs.CV2025-09被引 6

用完整放射科报告生成高保真3D胸部CT,提升临床细节还原度。

Radiology Report Conditional 3D CT Generation with Multi Encoder Latent diffusion Model

  • 融合三类医学文本编码器,全面捕捉报告中的临床语义。
  • 在CT RATE数据集上生成的3D CT体积与真实数据分布相似性达最优。
  • 适合需要高质量合成医学影像的研究者和医疗AI开发者。

文本到图像的潜空间扩散模型在医学图像合成中取得进展,但三维CT生成应用仍受限。现有方法依赖简化提示,忽略完整放射科报告中的丰富语义信息,导致图文对齐差、临床真实性不足。本文提出Report2CT框架,直接从自由文本放射科报告(含发现与印象部分)生成3D胸部CT体数据,采用多文本编码器机制。该模型整合三个预训练医学文本编码器(BiomedVLP CXR BERT、MedEmbed、ClinicalBERT),以捕获细微临床上下文。报告内容与体素间距信息共同条件化一个基于20000个CT体积的3D潜空间扩散模型(来自CT RATE数据集)。通过弗雷谢特图像距离(FID)评估生成数据分布相似性,使用CLIP相关指标衡量语义对齐,并与GenerateCT模型进行定量与定性对比。Report2CT生成具有解剖一致性、视觉质量优异且图文高度对齐的3D CT。多编码器条件化显著提升CLIP分数,表明更准确保留文本中的细粒度临床信息。无分类器引导进一步增强对齐,仅轻微降低FID表现。在MICCAI 2025 VLM3D挑战赛中,该模型在文本条件化CT生成任务上排名第一,并在所有评估指标上达到当前最优水平。通过利用完整放射科报告和多编码器文本条件化,Report2CT推动了3D CT合成发展,生成兼具临床真实性和高质量的合成数据。

原文摘要 · Abstract (English)

Text to image latent diffusion models have recently advanced medical image synthesis, but applications to 3D CT generation remain limited. Existing approaches rely on simplified prompts, neglecting the rich semantic detail in full radiology reports, which reduces text image alignment and clinical fidelity. We propose Report2CT, a radiology report conditional latent diffusion framework for synthesizing 3D chest CT volumes directly from free text radiology reports, incorporating both findings and impression sections using multiple text encoder. Report2CT integrates three pretrained medical text encoders (BiomedVLP CXR BERT, MedEmbed, and ClinicalBERT) to capture nuanced clinical context. Radiology reports and voxel spacing information condition a 3D latent diffusion model trained on 20000 CT volumes from the CT RATE dataset. Model performance was evaluated using Frechet Inception Distance (FID) for real synthetic distributional similarity and CLIP based metrics for semantic alignment, with additional qualitative and quantitative comparisons against GenerateCT model. Report2CT generated anatomically consistent CT volumes with excellent visual quality and text image alignment. Multi encoder conditioning improved CLIP scores, indicating stronger preservation of fine grained clinical details in the free text radiology reports. Classifier free guidance further enhanced alignment with only a minor trade off in FID. We ranked first in the VLM3D Challenge at MICCAI 2025 on Text Conditional CT Generation and achieved state of the art performance across all evaluation metrics. By leveraging complete radiology reports and multi encoder text conditioning, Report2CT advances 3D CT synthesis, producing clinically faithful and high quality synthetic data.

3D生成医学影像扩散模型文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。