arXiv:2503.22237cs.CV2025-03被引 2

融合SAM与CLIP优势,提升人体解析的语义与细节精度

SCHNet: SAM Marries CLIP for Human Parsing

  • 设计语义增强模块,融合CLIP语义特征与SAM空间细节
  • 提出高效微调模块,训练时间减少且性能显著提升
  • 在LIP/PPP/CIHP数据集上表现优异,适合高精度人体分割任务

视觉基础模型如分割一切模型(SAM)和对比语言-图像预训练模型(CLIP)在分割与检测任务中表现出色。尽管SAM擅长细粒度分割,但在语义感知分割中存在挑战;而CLIP虽具备强语义理解能力,但细粒度分割表现不足。人体解析需同时实现精确部件分割与高语义理解。基于SAM与CLIP的特性,本文提出高效模块以融合二者特征。设计语义增强模块,将CLIP的语义特征与SAM特征结合以提升解析效果;并提出高效微调模块,调整预训练SAM以适配人体解析任务,在保留空间细节的同时引入丰富语义信息,显著降低训练时间。大量实验表明,该方法在LIP、PPP和CIHP数据集上均取得显著效果。

原文摘要 · Abstract (English)

Vision Foundation Model (VFM) such as the Segment Anything Model (SAM) and Contrastive Language-Image Pre-training Model (CLIP) has shown promising performance for segmentation and detection tasks. However, although SAM excels in fine-grained segmentation, it faces major challenges when applying it to semantic-aware segmentation. While CLIP exhibits a strong semantic understanding capability via aligning the global features of language and vision, it has deficiencies in fine-grained segmentation tasks. Human parsing requires to segment human bodies into constituent parts and involves both accurate fine-grained segmentation and high semantic understanding of each part. Based on traits of SAM and CLIP, we formulate high efficient modules to effectively integrate features of them to benefit human parsing. We propose a Semantic-Refinement Module to integrate semantic features of CLIP with SAM features to benefit parsing. Moreover, we formulate a high efficient Fine-tuning Module to adjust the pretrained SAM for human parsing that needs high semantic information and simultaneously demands spatial details, which significantly reduces the training time compared with full-time training and achieves notable performance. Extensive experiments demonstrate the effectiveness of our method on LIP, PPP, and CIHP databases.

人体解析多模态融合模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。