利用医学影像的空间结构自监督学习,提升模型可解释性。
ISImed: A Framework for Self-Supervised Learning using Intrinsic Spatial Information in Medical Images
- 基于医学图像间结构相似性,构建位置编码的自监督目标
- 在两个公开数据集上实现优于主流SSL方法的表征性能
- 适合需要可解释医学表征的研究者与临床辅助诊断应用
本文证明,通过自监督学习(SSL)利用医学图像中的内在空间信息,可学习到可解释的表征。所提出的ISImed方法基于医学图像间人体结构变异极小的特点,利用多张图像间的结构一致性,建立一种自监督目标,使潜在表征能够捕捉图像区域的真实物理位置。具体而言,方法对图像块采样并计算所有可能组合的表示向量之间的距离矩阵,与真实空间距离进行对比。其核心思想是:学习到的潜在空间即为图像块的位置编码。我们假设,通过学习这些位置编码,将生成全面的图像表征。为验证该假设,我们在两个公开医学影像数据集上,将本方法与两种先进自监督学习基准方法进行了对比。结果表明,ISImed能高效学习到反映数据底层结构的表征,并可用于下游分类任务的迁移。
原文摘要 · Abstract (English)
This paper demonstrates that spatial information can be used to learn interpretable representations in medical images using Self-Supervised Learning (SSL). Our proposed method, ISImed, is based on the observation that medical images exhibit a much lower variability among different images compared to classic data vision benchmarks. By leveraging this resemblance of human body structures across multiple images, we establish a self-supervised objective that creates a latent representation capable of capturing its location in the physical realm. More specifically, our method involves sampling image crops and creating a distance matrix that compares the learned representation vectors of all possible combinations of these crops to the true distance between them. The intuition is, that the learned latent space is a positional encoding for a given image crop. We hypothesize, that by learning these positional encodings, comprehensive image representations have to be generated. To test this hypothesis and evaluate our method, we compare our learned representation with two state-of-the-art SSL benchmarking methods on two publicly available medical imaging datasets. We show that our method can efficiently learn representations that capture the underlying structure of the data and can be used to transfer to a downstream classification task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。