arXiv:2410.15744cs.CVcs.AI2024-10ICLR被引 8

通过视觉-语言对齐实现3D医学图像零样本病灶分割,提升未见病灶识别能力。

Unleashing the Potential of Vision-Language Pre-Training for 3D Zero-Shot Lesion Segmentation via Mask-Attribute Alignment

  • 设计多尺度掩码-属性对齐框架,显式关联病灶视觉特征与语义描述
  • 在3个数据集12类病灶上实现优于现有方法的零样本分割性能
  • 适合医疗影像分析、跨模态学习研究者关注,尤其关注零样本场景

近年来,医学视觉-语言预训练模型在零样本疾病识别方面取得显著进展。然而,将图像级知识迁移至像素级任务(如3D CT扫描中的病灶分割)仍是关键挑战。由于病理视觉特征复杂多变,现有方法难以将训练中未遇到的细粒度病灶特征与疾病相关文本表示对齐。本文提出Malenia,一种专为3D零样本病灶分割设计的多尺度病变级掩码-属性对齐框架。Malenia提升掩码表征与其基本属性之间的兼容性,明确关联未见病灶的视觉特征与先前学习的可扩展知识。此外,我们设计了跨模态知识注入模块,双向增强视觉与文本特征,有效引导分割结果生成。在三个数据集和12类病灶上的全面实验验证了Malenia的优越性能。

原文摘要 · Abstract (English)

Recent advancements in medical vision-language pre-training models have driven significant progress in zero-shot disease recognition. However, transferring image-level knowledge to pixel-level tasks, such as lesion segmentation in 3D CT scans, remains a critical challenge. Due to the complexity and variability of pathological visual characteristics, existing methods struggle to align fine-grained lesion features not encountered during training with disease-related textual representations. In this paper, we present Malenia, a novel multi-scale lesion-level mask-attribute alignment framework, specifically designed for 3D zero-shot lesion segmentation. Malenia improves the compatibility between mask representations and their associated elemental attributes, explicitly linking the visual features of unseen lesions with the extensible knowledge learned from previously seen ones. Furthermore, we design a Cross-Modal Knowledge Injection module to enhance both visual and textual features with mutually beneficial information, effectively guiding the generation of segmentation results. Comprehensive experiments across three datasets and 12 lesion categories validate the superior performance of Malenia.

3D分割零样本学习视觉-语言对齐医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。