用多尺度向量量化VAE生成胶囊内镜图像,解决医学数据稀缺问题。
Multiscale Vector-Quantized Variational Autoencoder for Endoscopic Image Synthesis
- 提出多尺度向量量化变分自编码器,提升图像生成质量与多样性。
- 可条件生成包含不同类型异常的合成图像,性能接近真实数据训练模型。
- 适用于多种胃肠道异常生成,适合医疗AI数据增强场景。
无线胶囊内镜(WCE)产生大量图像需人工筛查。深度学习临床决策支持(CDS)系统可辅助筛查,但其性能依赖大规模多样数据集。由于隐私限制和标注成本,此类数据稀缺,阻碍了CDS发展。生成式机器学习为此提供解决方案。现有生成方法如GAN和VAE常面临训练不稳定、视觉多样性不足等问题,尤其在生成异常图像时表现不佳。本文提出一种新型基于VAE的医学图像合成方法,并应用于WCE图像生成。主要贡献包括:a) 提出多尺度向量量化变分自编码器(MSVQ-VAE);b) 首次将异常无缝融入正常WCE图像;c) 实现条件生成,可引入不同类型的异常;d) 在多种异常类型(息肉、血管病变、炎症)上验证。通过图像分类评估生成图像对CDS的效用,对比实验表明:使用本方法生成的异常图像训练的分类器,性能与仅用真实数据训练的模型相当。该方法具有广泛适用性,可推广至其他医学多媒体领域。
原文摘要 · Abstract (English)
Gastrointestinal (GI) imaging via Wireless Capsule Endoscopy (WCE) generates a large number of images requiring manual screening. Deep learning-based Clinical Decision Support (CDS) systems can assist screening, yet their performance relies on the existence of large, diverse, training medical datasets. However, the scarcity of such data, due to privacy constraints and annotation costs, hinders CDS development. Generative machine learning offers a viable solution to combat this limitation. While current Synthetic Data Generation (SDG) methods, such as Generative Adversarial Networks and Variational Autoencoders have been explored, they often face challenges with training stability and capturing sufficient visual diversity, especially when synthesizing abnormal findings. This work introduces a novel VAE-based methodology for medical image synthesis and presents its application for the generation of WCE images. The novel contributions of this work include a) multiscale extension of the Vector Quantized VAE model, named as Multiscale Vector Quantized Variational Autoencoder (MSVQ-VAE); b) unlike other VAE-based SDG models for WCE image generation, MSVQ-VAE is used to seamlessly introduce abnormalities into normal WCE images; c) it enables conditional generation of synthetic images, enabling the introduction of different types of abnormalities into the normal WCE images; d) it performs experiments with a variety of abnormality types, including polyps, vascular and inflammatory conditions. The utility of the generated images for CDS is assessed via image classification. Comparative experiments demonstrate that training a CDS classifier using the abnormal images generated by the proposed methodology yield comparable results with a classifier trained with only real data. The generality of the proposed methodology promises its applicability to various domains related to medical multimedia.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。