通过解耦视觉语言属性提升跨域人物检索的持续学习能力
Vision-Language Attribute Disentanglement and Reinforcement for Lifelong Person Re-Identification
- 显式解耦全局与局部文本属性,增强细粒度特征提取
- 跨模态与跨域对齐使新旧知识协同,遗忘率降低2.1%-2.5%
- 适合需要长期更新、多场景泛化的行人识别系统
终身人物再识别(LReID)旨在从不同数据域中持续学习,构建统一的行人检索模型。现有方法多从零开始或基于图像分类预训练模型,而视觉-语言模型(VLM)在多种任务中展现出强泛化能力。尽管可直接适配至VLM,但现有方法仅关注全局特征学习,忽视细粒度属性知识,导致知识获取与抗遗忘能力受限。为此,本文提出一种新型VLM驱动的LReID方法——视觉语言属性解耦与强化(VLADR)。核心思想是显式建模通用人体属性,促进跨域知识迁移,从而有效利用历史知识强化新知识学习并缓解遗忘。具体而言,VLADR包含多粒度文本属性解耦机制,挖掘图像的全局与多样化局部文本属性;设计跨域跨模态属性强化方案,通过跨模态对齐引导视觉属性提取,并利用跨域对齐实现细粒度知识迁移。实验表明,相比当前最优方法,本方法在抗遗忘与泛化能力上分别提升1.9%-2.2%和2.1%-2.5%。源代码已开源。
原文摘要 · Abstract (English)
Lifelong person re-identification (LReID) aims to learn from varying domains to obtain a unified person retrieval model. Existing LReID approaches typically focus on learning from scratch or a visual classification-pretrained model, while the Vision-Language Model (VLM) has shown generalizable knowledge in a variety of tasks. Although existing methods can be directly adapted to the VLM, since they only consider global-aware learning, the fine-grained attribute knowledge is underleveraged, leading to limited acquisition and anti-forgetting capacity. To address this problem, we introduce a novel VLM-driven LReID approach named Vision-Language Attribute Disentanglement and Reinforcement (VLADR). Our key idea is to explicitly model the universally shared human attributes to improve inter-domain knowledge transfer, thereby effectively utilizing historical knowledge to reinforce new knowledge learning and alleviate forgetting. Specifically, VLADR includes a Multi-grain Text Attribute Disentanglement mechanism that mines the global and diverse local text attributes of an image. Then, an Inter-domain Cross-modal Attribute Reinforcement scheme is developed, which introduces cross-modal attribute alignment to guide visual attribute extraction and adopts inter-domain attribute alignment to achieve fine-grained knowledge transfer. Experimental results demonstrate that our VLADR outperforms the state-of-the-art methods by 1.9\%-2.2\% and 2.1\%-2.5\% on anti-forgetting and generalization capacity. Our source code is available at https://github.com/zhoujiahuan1991/CVPR2026-VLADR
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。