arXiv:2506.10230eess.IVcs.CV2025-06中稿 · publication in IEE…被引 2

用临床报告和分类标签生成高质量前列腺MRI,数据少也能用。

Leveraging Clinical Text and Class Conditioning for 3D Prostate MRI Generation

  • 双头条件控制:同时用病历文本和诊断分类训练扩散模型。
  • 3D FID仅0.025,优于现有模型(0.070),合成图像质量高。
  • 用生成图像训练分类器,准确率从69%提升至74%,适合小样本研究。

目标:潜在扩散模型(LDM)可缓解医学影像机器学习中的数据稀缺问题。但现有医学LDM方法通常依赖短提示文本编码器、非医学LDM或大量数据,限制了性能与科学可及性。本文提出一种新型LDM条件化方法。方法:提出类条件高效大语言模型适配器(CCELLA),一种双头条件化策略,同时以自由文本临床报告和放射科分类信息条件化LDM U-Net。并设计基于CCELLA的数据高效LDM流水线及联合损失函数。在3D前列腺MRI上评估性能,对比先进方法;随后将生成图像用于下游分类器训练数据增强。结果:在有限规模的3D前列腺MRI数据集上,本方法取得0.025的3D FID,显著优于近期基础模型(FID 0.070)。在前列腺癌预测分类器训练中,加入生成图像后准确率由69%提升至74%,优于先前最先进方法。仅使用本方法生成图像训练的分类器,性能接近真实图像训练。结论:本方法在数据量少、人工标注极少的情况下,提升了合成图像质量与下游分类性能。意义:该以CCELLA为核心的流程,实现了报告与类别条件化的高质医学图像生成,适用于低数据量与低标注成本场景,提升LDM性能与科研可及性。

原文摘要 · Abstract (English)

Objective: Latent diffusion models (LDM) could alleviate data scarcity challenges affecting machine learning development for medical imaging. However, medical LDM strategies typically rely on short-prompt text encoders, nonmedical LDMs, or large data volumes. These strategies can limit performance and scientific accessibility. We propose a novel LDM conditioning approach to address these limitations. Methods: We propose Class-Conditioned Efficient Large Language model Adapter (CCELLA), a novel dual-head conditioning approach that simultaneously conditions the LDM U-Net with free-text clinical reports and radiology classification. We also propose a data-efficient LDM pipeline centered around CCELLA and a proposed joint loss function. We first evaluate our method on 3D prostate MRI against state-of-the-art. We then augment a downstream classifier model training dataset with synthetic images from our method. Results: Our method achieves a 3D FID score of 0.025 on a size-limited 3D prostate MRI dataset, significantly outperforming a recent foundation model with FID 0.070. When training a classifier for prostate cancer prediction, adding synthetic images generated by our method during training improves classifier accuracy from 69% to 74% and outperforms classifiers trained on images generated by prior state-of-the-art. Classifier training solely on our method's synthetic images achieved comparable performance to real image training. Conclusion: We show that our method improved both synthetic image quality and downstream classifier performance using limited data and minimal human annotation. Significance: The proposed CCELLA-centric pipeline enables radiology report and class-conditioned LDM training for high-quality medical image synthesis given limited data volume and human data annotation, improving LDM performance and scientific accessibility.

医学图像生成扩散模型小样本学习前列腺MRI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。