arXiv:2411.12201cs.CV2024-11被引 7

提出形状不变表示学习,提升图像分类在不同环境下的鲁棒性。

Invariant Shape Representation Learning For Image Classification

  • 通过可变形变换的潜在形状空间联合学习不变特征。
  • 在2D模拟图、3D脑部和心血管MRI上均实现更准确的分类结果。
  • 适合需要跨环境稳定预测的医学图像分析场景。

几何形状特征广泛用作图像分类的强预测因子。然而,现有分类器(如深度神经网络)通常直接利用这些形状特征与目标变量之间的统计相关性,而此类相关性在不同环境(如不同年龄组)中可能具有伪相关性和不稳定性,导致预测偏差或不准确。本文首次提出不变形状表示学习(ISRL)框架,以增强图像分类器的鲁棒性。与主流方法仅在图像空间提取特征不同,ISRL模型通过可变形变换参数化的潜在形状空间,联合捕捉跨环境的不变特征。为此,我们基于不变风险最小化(IRM)构建新学习范式,在多个训练分布/环境中学习图像与形状特征的不变表示。通过嵌入对目标变量在不同环境中保持不变的特征,模型实现了更一致且准确的预测。我们在2D模拟图像、真实3D脑部及心动电影磁共振成像(cine cardiovascular MRI)上验证了该方法的有效性。代码已公开于https://github.com/tonmoy-hossain/ISRL。

原文摘要 · Abstract (English)

Geometric shape features have been widely used as strong predictors for image classification. Nevertheless, most existing classifiers such as deep neural networks (DNNs) directly leverage the statistical correlations between these shape features and target variables. However, these correlations can often be spurious and unstable across different environments (e.g., in different age groups, certain types of brain changes have unstable relations with neurodegenerative disease); hence leading to biased or inaccurate predictions. In this paper, we introduce a novel framework that for the first time develops invariant shape representation learning (ISRL) to further strengthen the robustness of image classifiers. In contrast to existing approaches that mainly derive features in the image space, our model ISRL is designed to jointly capture invariant features in latent shape spaces parameterized by deformable transformations. To achieve this goal, we develop a new learning paradigm based on invariant risk minimization (IRM) to learn invariant representations of image and shape features across multiple training distributions/environments. By embedding the features that are invariant with regard to target variables in different environments, our model consistently offers more accurate predictions. We validate our method by performing classification tasks on both simulated 2D images, real 3D brain and cine cardiovascular magnetic resonance images (MRIs). Our code is publicly available at https://github.com/tonmoy-hossain/ISRL.

形状表示不变学习医学图像鲁棒分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。