提出SDICE指标,量化合成医学图像多样性。
Introducing SDICE: An Index for Assessing Diversity of Synthetic Medical Datasets
- 用预训练对比编码器计算图像相似度分布
- 通过分布距离衡量合成数据与真实数据差异
- 适用于评估医学图像生成质量,适合研究者使用
生成模型在合成医学图像方面取得进展,可作为数据增强提升医疗图像分析模型性能。尽管合成图像保真度不断提升,其多样性仍缺乏系统评估。本文提出SDICE指数,基于对比编码器对真实与合成图像的相似度分布进行建模,通过计算两者分布间的距离,并经指数函数归一化,得到跨领域可比的统一度量。在MIMIC-chest X-ray和ImageNet数据集上的实验验证了该指标的有效性。
原文摘要 · Abstract (English)
Advancements in generative modeling are pushing the state-of-the-art in synthetic medical image generation. These synthetic images can serve as an effective data augmentation method to aid the development of more accurate machine learning models for medical image analysis. While the fidelity of these synthetic images has progressively increased, the diversity of these images is an understudied phenomenon. In this work, we propose the SDICE index, which is based on the characterization of similarity distributions induced by a contrastive encoder. Given a synthetic dataset and a reference dataset of real images, the SDICE index measures the distance between the similarity score distributions of original and synthetic images, where the similarity scores are estimated using a pre-trained contrastive encoder. This distance is then normalized using an exponential function to provide a consistent metric that can be easily compared across domains. Experiments conducted on the MIMIC-chest X-ray and ImageNet datasets demonstrate the effectiveness of SDICE index in assessing synthetic medical dataset diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。