用生成模型合成3D前列腺MRI,提升小数据机构的癌症检测性能
Mitigating 3D Prostate Biparametric MRI Data Scarcity through Domain Adaptation using Locally-Trained Latent Diffusion Models for Prostate Cancer Detection
- 基于局部训练的潜在扩散模型,同步生成T2、ADC和高b值影像
- 合成数据预训练使分类器在小样本外数据上准确率提升12.5%以上
- 适合医疗数据稀缺机构,推动医学影像AI跨中心应用
目的:潜在扩散模型(LDM)可缓解医学图像分析中数据稀缺问题。近期的CCELLA LDM通过合成MRI提升前列腺癌检测性能,但仅限于轴向T2加权(AxT2)序列,未研究机构间域偏移,且侧重PI-RADS而非组织病理学结果。方法:本文提出CCELLA++,一种新型LDM流程,可同步生成3D双参数前列腺MRI(bpMRI),包括AxT2、高b值扩散序列(HighB)和表观扩散系数图(ADC),以克服上述局限。我们采用无源域自适应策略,先用单机构真实或合成数据预训练分类器,再在外部分布外数据的少量样本上微调。结果:CCELLA++在AxT2上的核 inception 距离与CCELLA相当(0.0128 vs 0.0131)。CCELLA++合成bpMRI预训练在外部数据集体积≤166时,平均精度(AP)和曲线下面积(AUC)提升最高达12.5%(p<0.01),无预训练情况下在外部数据量为332时,AUC提升达25%(p<0.05),且优于仅用CCELLA AxT2生成数据预训练,在小样本(n=83,p<0.001 AP和AUC)和全数据(n=1329,p<0.05 AP和AUC)场景下均表现更优。结论:CCELLA++生成的合成bpMRI可显著提升下游分类器的泛化能力与性能,优于真实bpMRI或仅用AxT2生成的数据。未来工作应量化医学图像质量,平衡bpMRI LDM训练,并引入额外条件信息。意义:CCELLA++可生成超越真实数据的合成bpMRI,适用于数据稀缺的外部机构域自适应,推动医学影像机器学习发展。代码已开源:https://github.com/grabkeem/CCELLA-plus-plus
原文摘要 · Abstract (English)
Objective: Latent diffusion models (LDMs) could mitigate data scarcity challenges affecting machine learning development for medical image interpretation. The recent CCELLA LDM improved prostate cancer detection performance using synthetic MRI for classifier training but was limited to the axial T2-weighted (AxT2) sequence, did not investigate inter-institutional domain shift, and prioritized PI-RADS over histopathology outcomes. Methods: We propose CCELLA++, a novel LDM pipeline for simultaneous 3D biparametric prostate MRI (bpMRI) generation, including the AxT2, high b-value diffusion series (HighB) and apparent diffusion coefficient map (ADC), to overcome these limitations. We investigated source-free domain adaptation with classifiers pretrained on single institution real or LDM-generated synthetic data prior to fine-tuning on fractions of an out-of-distribution, external dataset. Results: CCELLA++ achieved comparable AxT2 Kernel Inception Distance to CCELLA (0.0128, 0.0131 respectively). CCELLA++ synthetic bpMRI pretraining outperformed real bpMRI in AP and AUC up to 12.5% (n<=166) external dataset volume (p<0.01 all), no pretraining in AUC up to 25% external volume (n=332, p<0.05 all), and CCELLA AxT2-only pretraining in both data-scarce (n=83, p<0.001 AP and AUC) and full data (n=1329, p<0.05 AP and AUC) scenarios. Conclusion: CCELLA++ synthetic bpMRI can improve downstream classifier generalization and performance beyond real bpMRI or CCELLA-generated AxT2-only images. Future work should quantify medical image quality, balance bpMRI LDM training, and condition the LDM with additional information. Significance: CCELLA++ can generate synthetic bpMRI that outperforms real data for domain adaptation with data-scarce external institutions, advancing machine learning development for medical imaging. Our code is available at https://github.com/grabkeem/CCELLA-plus-plus
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。