arXiv:2606.28980cs.CVcs.AI2026-06

用临床数据生成卵巢癌3D CT影像,解决报告缺失难题

Evidence-Based Text-Conditioned 3D CT Synthesis for Ovarian Cancer

论文配图:Evidence-Based Text-Conditioned 3D CT Synthesis for Ovarian Cancer
图 1 · 摘自论文原文
  • 基于影像特征和临床数据自动生成报告文本,不依赖原始报告
  • 生成的3D CT在解剖结构和强度上与真实数据接近(FID2.5D 29.35)
  • 适用于缺乏报告标注的卵巢癌影像研究,推动合成数据应用

卵巢癌常于晚期确诊,术前增强CT对分期与手术规划至关重要;但标注影像数据稀缺且受隐私限制,制约了通用模型的发展。现有文本条件3D CT生成方法依赖配对报告,且仅在胸部CT上验证。本文提出OvESyn框架,通过CT影像描述符与常规临床数据构建标准化“发现”与“印象”文本,无需原始报告,并将其用于条件化适配493例高级别浆液性卵巢癌患者的潜空间扩散模型。这是首个针对腹盆腔肿瘤场景的文本条件3D CT生成框架。系统消融实验表明,生成器领域适应是跨越域差距的关键机制;若无此步骤,生成结果仍锚定于胸腔预训练域,精度与召回降至零,FID2.5D超140;而编码器对齐则优化强度与细节。完整版OvESyn实现最优分布与强度保真度(FID2.5D 29.35,Precision 0.671,Wasserstein-1 0.044),仅生成器微调版本覆盖更广(Recall 0.645),体现保真度与覆盖率的权衡由编码器适配决定。仅需自动分割与术前常规数据,该框架支持在报告稀缺环境下迁移应用,为腹盆腔肿瘤影像合成数据集构建奠定基础。

原文摘要 · Abstract (English)

Ovarian cancer is frequently diagnosed at an advanced stage, making preoperative contrast-enhanced computed tomography (CT) central to staging and surgical planning; yet the scarcity of annotated imaging data, compounded by privacy regulations, limits the development of generalizable computational models in this domain. Text-conditioned 3D CT synthesis has shown promise, but existing pipelines depend on paired radiology reports and have been evaluated only on chest CT. We propose OvESyn (Ovarian Evidence-based Synthesis), a framework that constructs standardized Findings and Impression sections directly from CT-derived imaging descriptors and routine clinical metadata, without any original radiology report, and uses them to condition a latent diffusion model adapted to 493 high-grade serous ovarian carcinoma patients. This is the first text-conditioned 3D CT synthesis framework adapted to an abdomino-pelvic oncologic setting. A systematic ablation over two adaptation axes, vision-language encoder alignment and generator fine-tuning, identifies generator domain adaptation as the operative mechanism for crossing the domain gap and establishing the target anatomy: without it, synthesis remains anchored to the thoracic pretraining domain, with Precision and Recall collapsing to zero and FID2.5D exceeding 140, regardless of encoder alignment. Encoder alignment instead refines intensity and fine detail. The full OvESyn attains the best distributional and intensity fidelity (FID2.5D 29.35, Precision 0.671, Wasserstein-1 0.044), while the generator-only variant maximizes coverage (Recall 0.645), reflecting a fidelity/coverage trade-off governed by encoder adaptation. Requiring only automatic segmentations and routine preoperative metadata, OvESyn supports transferability to report-scarce settings and provides a foundation for synthetic cohort generation in abdomino-pelvic oncologic imaging.

3D生成卵巢癌合成数据扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。