用风格迁移增强乳腺钼靶图像数据多样性,提升模型泛化能力
Style transfer as data augmentation: evaluating unpaired image-to-image translation models in mammography
- 通过无配对图像翻译模型迁移不同数据集的图像风格
- 对比CycleGAN与SynDiff在三个乳腺钼靶数据集上的表现
- 强调需多指标评估,避免单一指标误导
深度学习模型可从乳腺钼靶图像中检测乳腺癌,但过拟合和泛化能力差限制了其临床应用。不同人群间因扫描技术或患者特征差异导致数据分布不同,使模型性能下降。数据增强可通过改变现有样本的特征表示来提升多样性。图像到图像翻译模型能将一个数据集的特征风格迁移到另一个。然而,医学影像缺乏真实标签,模型评估困难。本文分析了评估风格迁移算法的关键因素,比较了主流评估指标的优劣,并以两个生成模型(CycleGAN与SynDiff)在三个乳腺钼靶数据集上进行无配对图像翻译。研究指出某些模型缺陷会影响特定指标有效性,并揭示不同指标衡量的是模型性能的不同方面,强调应结合多种指标进行全面评估。
原文摘要 · Abstract (English)
Several studies indicate that deep learning models can learn to detect breast cancer from mammograms (X-ray images of the breasts). However, challenges with overfitting and poor generalisability prevent their routine use in the clinic. Models trained on data from one patient population may not perform well on another due to differences in their data domains, emerging due to variations in scanning technology or patient characteristics. Data augmentation techniques can be used to improve generalisability by expanding the diversity of feature representations in the training data by altering existing examples. Image-to-image translation models are one approach capable of imposing the characteristic feature representations (i.e. style) of images from one dataset onto another. However, evaluating model performance is non-trivial, particularly in the absence of ground truths (a common reality in medical imaging). Here, we describe some key aspects that should be considered when evaluating style transfer algorithms, highlighting the advantages and disadvantages of popular metrics, and important factors to be mindful of when implementing them in practice. We consider two types of generative models: a cycle-consistent generative adversarial network (CycleGAN) and a diffusion-based SynDiff model. We learn unpaired image-to-image translation across three mammography datasets. We highlight that undesirable aspects of model performance may determine the suitability of some metrics, and also provide some analysis indicating the extent to which various metrics assess unique aspects of model performance. We emphasise the need to use several metrics for a comprehensive assessment of model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。