用生成模型合成医学影像,验证其在肿瘤和骨骼分割中的实用性。
Enhancing Privacy: The Utility of Stand-Alone Synthetic CT and MRI for Tumor and Bone Segmentation
- 用GAN和扩散模型生成头颈CT与脑胶质瘤MRI数据。
- 合成MRI分割效果好(DSC=0.834),CT效果差(DSC=0.064)。
- 合成数据可独立用于分割,尤其适合结构简单的任务。
人工智能需要大量数据,但医疗数据受严格隐私保护。匿名化虽必要,但在头颈部等结构重叠区域存在挑战。合成数据或可解决此问题,但现有研究常缺乏对真实性和实用性的严格评估。本文研究合成数据在分割任务中替代真实数据的程度。使用来自两个大型数据集的头颈癌CT与脑胶质瘤MRI数据,通过生成对抗网络和扩散模型生成合成数据。采用MAE、MS-SSIM、放射组学及5名放射科医生参与的视觉图灵测试(VTT)评估合成数据质量,以骰子系数(DSC)评估分割实用性。放射组学显示合成MRI保真度高,但合成CT组织真实性较差,肿瘤相关性系数分别为0.8784(MRI)与0.5461(CT)。DSC结果显示:肿瘤分割在CT上仅达0.064,在MRI上为0.834;骨分割平均DSC为0.841。观察到DSC与相关性之间存在关联,但受限于任务复杂性。视觉图灵测试表明合成CT具有一定实用性,但教育应用有限。尽管合成数据可独立用于分割任务,但仍受目标结构复杂性制约。提升生成模型对异质输入的鲁棒性并学习细微特征,是增强其真实感与拓展应用的关键。
原文摘要 · Abstract (English)
AI requires extensive datasets, while medical data is subject to high data protection. Anonymization is essential, but poses a challenge for some regions, such as the head, as identifying structures overlap with regions of clinical interest. Synthetic data offers a potential solution, but studies often lack rigorous evaluation of realism and utility. Therefore, we investigate to what extent synthetic data can replace real data in segmentation tasks. We employed head and neck cancer CT scans and brain glioma MRI scans from two large datasets. Synthetic data were generated using generative adversarial networks and diffusion models. We evaluated the quality of the synthetic data using MAE, MS-SSIM, Radiomics and a Visual Turing Test (VTT) performed by 5 radiologists and their usefulness in segmentation tasks using DSC. Radiomics indicates high fidelity of synthetic MRIs, but fall short in producing highly realistic CT tissue, with correlation coefficient of 0.8784 and 0.5461 for MRI and CT tumors, respectively. DSC results indicate limited utility of synthetic data: tumor segmentation achieved DSC=0.064 on CT and 0.834 on MRI, while bone segmentation a mean DSC=0.841. Relation between DSC and correlation is observed, but is limited by the complexity of the task. VTT results show synthetic CTs' utility, but with limited educational applications. Synthetic data can be used independently for the segmentation task, although limited by the complexity of the structures to segment. Advancing generative models to better tolerate heterogeneous inputs and learn subtle details is essential for enhancing their realism and expanding their application potential.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。