用扩散模型生成牙科全景片,再超分辨率提升细节。
PanoDiff-SR: Synthesizing Dental Panoramic Radiographs using Diffusion and Super-resolution
- 先用扩散模型生成低分辨率牙科全景片,再通过新架构超分辨率重建。
- 生成图像与真实图像的相似度达40.69(弗雷歇距离),专家仅能辨识出68.5%。
- 适合医学影像生成、人工智能训练及教学使用。
近年来,合成高质量、逼真的医疗图像受到广泛关注。此类合成数据集可缓解人工智能研究中公开数据集稀缺的问题,并可用于教育目的。本文提出一种结合扩散生成(PanoDiff)与超分辨率(SR)的方法,用于生成合成牙科全景放射图像(PR)。前者生成低分辨率(LR)种子图像(256×128),再由超分辨率模型处理得到高分辨率(HR)图像(1024×512)。针对超分辨率,我们提出一种新型基于变压器的架构,有效学习局部-全局关系,显著提升边缘与纹理清晰度。实验结果表明,7243张真实与合成图像之间的弗雷歇起始距离为40.69。真实与合成高分辨率及低分辨率图像的Inception分数分别为2.55、2.30、2.90和2.98。六名临床专家在限定时间内评估100张合成与100张真实图像混合样本,平均辨别准确率为68.5%(随机猜测为50%)。
原文摘要 · Abstract (English)
There has been increasing interest in the generation of high-quality, realistic synthetic medical images in recent years. Such synthetic datasets can mitigate the scarcity of public datasets for artificial intelligence research, and can also be used for educational purposes. In this paper, we propose a combination of diffusion-based generation (PanoDiff) and Super-Resolution (SR) for generating synthetic dental panoramic radiographs (PRs). The former generates a low-resolution (LR) seed of a PR (256 X 128) which is then processed by the SR model to yield a high-resolution (HR) PR of size 1024 X 512. For SR, we propose a state-of-the-art transformer that learns local-global relationships, resulting in sharper edges and textures. Experimental results demonstrate a Frechet inception distance score of 40.69 between 7243 real and synthetic images (in HR). Inception scores were 2.55, 2.30, 2.90 and 2.98 for real HR, synthetic HR, real LR and synthetic LR images, respectively. Among a diverse group of six clinical experts, all evaluating a mixture of 100 synthetic and 100 real PRs in a time-limited observation, the average accuracy in distinguishing real from synthetic images was 68.5% (with 50% corresponding to random guessing).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。