用语言引导3D CT图像分割,自动聚焦难分区域。
VOILA: Complexity-Aware Universal Segmentation of CT images by Voxel Interacting with Language
- 通过体素与文本对齐,基于相似度分类
- 复杂度感知采样提升难分区域精度
- 少参数低计算,跨数据集无需微调
近期在CT图像通用分割上取得显著进展。受视觉-语言方法启发,越来越多研究采用文本提示和对比学习构建通用分割模型。然而,3D图像与文本提示间存在显著的信息密度差异。此外,标准全连接层分割方法在多类别场景下表现不佳,泛化能力有限。为此,我们提出体素与语言交互方法(VOILA),首先将体素与语言映射到共享表示空间,并基于余弦相似度进行体素分类;随后设计体素-语言交互框架,缓解因前景-背景差异及目标体积变化导致的类别不平衡问题;进一步提出复杂度感知采样策略,通过可训练高斯混合分布生成伪热图,聚焦于难以分割的区域。实验表明,所提方法在减少参数量与训练计算成本的同时,实现了更优性能,并在多个数据集间表现出显著泛化能力,无需额外微调。
原文摘要 · Abstract (English)
Satisfactory progress has been achieved recently in universal segmentation of CT images. Following the success of vision-language methods, there is a growing trend towards utilizing text prompts and contrastive learning to develop universal segmentation models. However, there exists a significant imbalance in information density between 3D images and text prompts. Moreover, the standard fully connected layer segmentation approach faces significant challenges in handling multiple classes and exhibits poor generalizability. To address these challenges, we propose the VOxel Interacting with LAnguage method (VOILA) for universal CT image segmentation. Initially, we align voxels and language into a shared representation space and classify voxels on the basis of cosine similarity. Subsequently, we develop the Voxel-Language Interaction framework to mitigate the impact of class imbalance caused by foreground-background discrepancies and variations in target volumes. Furthermore, a Complexity-Aware Sampling method is proposed to focus on region hard to segment, achieved by generating pseudo-heatmaps from a trainable Gaussian mixture distribution. Our results indicate the proposed VOILA is capable to achieve improved performance with reduced parameters and computational cost during training. Furthermore, it demonstrates significant generalizability across diverse datasets without additional fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。