用CLIP生成多粒度语义标签,提升人像视觉任务的自监督预训练效果
CLIP-Guided Adaptable Self-Supervised Learning for Human-Centric Visual Tasks
- 利用CLIP生成身体部位和属性等多层级伪标签
- 在多个基准上超越现有自监督方法,提升人像分析性能
- 支持不同下游任务动态适配,适合多样化人像视觉场景
人像视觉分析在监控、医疗和人机交互等领域至关重要。随着大规模无标注人像数据集的出现,亟需一种通用的无监督预训练模型以支持多样化的下游任务。为此,我们提出CLASP(CLIP-guided Adaptable Self-suPervised learning),一种面向人像视觉任务的新型无监督预训练框架。CLASP利用强大的视觉-语言模型CLIP生成低层(如身体部位)和高层(如属性)语义伪标签,并将其融入视觉表征,增强表达能力与泛化性。针对不同下游任务对语义粒度需求不同,CLASP引入提示控制的专家混合(Prompt-Controlled MoE)模块,动态调整特征提取以缓解特征冲突并提升可迁移性。此外,采用多任务预训练策略,由CLIP生成的部件与属性级伪标签引导表征学习。在多个基准上的大量实验表明,CLASP持续优于现有自监督预训练方法,推动了人像视觉分析的发展。
原文摘要 · Abstract (English)
Human-centric visual analysis plays a pivotal role in diverse applications, including surveillance, healthcare, and human-computer interaction. With the emergence of large-scale unlabeled human image datasets, there is an increasing need for a general unsupervised pre-training model capable of supporting diverse human-centric downstream tasks. To achieve this goal, we propose CLASP (CLIP-guided Adaptable Self-suPervised learning), a novel framework designed for unsupervised pre-training in human-centric visual tasks. CLASP leverages the powerful vision-language model CLIP to generate both low-level (e.g., body parts) and high-level (e.g., attributes) semantic pseudo-labels. These multi-level semantic cues are then integrated into the learned visual representations, enriching their expressiveness and generalizability. Recognizing that different downstream tasks demand varying levels of semantic granularity, CLASP incorporates a Prompt-Controlled Mixture-of-Experts (MoE) module. MoE dynamically adapts feature extraction based on task-specific prompts, mitigating potential feature conflicts and enhancing transferability. Furthermore, CLASP employs a multi-task pre-training strategy, where part- and attribute-level pseudo-labels derived from CLIP guide the representation learning process. Extensive experiments across multiple benchmarks demonstrate that CLASP consistently outperforms existing unsupervised pre-training methods, advancing the field of human-centric visual analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。