让医学影像自动拆解出器官结构与风格,提升模型可解释性。
ConceptVAE: Self-Supervised Fine-Grained Concept Disentanglement from 2D Echocardiographies
- 通过自监督学习将图像分解为固定数量的细粒度概念与局部风格。
- 在心超图上成功识别血池、室间隔等结构,多项任务性能超越传统方法。
- 适合需要可解释性的医疗影像分析场景,如异常检测与合成数据生成。
尽管传统自监督学习在多种医疗任务中提升了性能与鲁棒性,但其依赖单向量嵌入,难以捕捉如解剖结构或器官等细粒度概念。无需标注即可识别此类概念及其特征,有望改进预训练方法,并推动细粒度图像检索与基于概念的异常检测等新应用。本文提出 ConceptVAE,一种新颖的预训练框架,可在自监督条件下检测并解耦2D心脏超声图像中的细粒度概念与其风格特征。设计了一套损失项与模型结构原语,将输入数据离散化为预设数量的概念及其局部风格。定性和定量验证表明,ConceptVAE能有效识别血池、室间隔壁等细粒度解剖结构。量化结果显示,其在基于区域的实例检索、语义分割、分布外检测和目标检测任务中均优于传统自监督方法。此外,我们探索了生成保持相同概念但风格不同的分布内合成数据,展现了更精准的数据生成潜力。整体而言,本研究引入并验证了一种基于概念-风格解耦的有前景的预训练技术,为开发更具可解释性的医学图像分析模型开辟了新路径。
原文摘要 · Abstract (English)
While traditional self-supervised learning methods improve performance and robustness across various medical tasks, they rely on single-vector embeddings that may not capture fine-grained concepts such as anatomical structures or organs. The ability to identify such concepts and their characteristics without supervision has the potential to improve pre-training methods, and enable novel applications such as fine-grained image retrieval and concept-based outlier detection. In this paper, we introduce ConceptVAE, a novel pre-training framework that detects and disentangles fine-grained concepts from their style characteristics in a self-supervised manner. We present a suite of loss terms and model architecture primitives designed to discretise input data into a preset number of concepts along with their local style. We validate ConceptVAE both qualitatively and quantitatively, demonstrating its ability to detect fine-grained anatomical structures such as blood pools and septum walls from 2D cardiac echocardiographies. Quantitatively, ConceptVAE outperforms traditional self-supervised methods in tasks such as region-based instance retrieval, semantic segmentation, out-of-distribution detection, and object detection. Additionally, we explore the generation of in-distribution synthetic data that maintains the same concepts as the training data but with distinct styles, highlighting its potential for more calibrated data generation. Overall, our study introduces and validates a promising new pre-training technique based on concept-style disentanglement, opening multiple avenues for developing models for medical image analysis that are more interpretable and explainable than black-box approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。