无监督跨数据集图文行人检索新方法,提升实际应用泛化能力。
Graph-Based Cross-Domain Knowledge Distillation for Cross-Dataset Text-to-Image Person Retrieval
- 构建图结构传播跨模态特征,建模跨域关联
- 对比动量知识蒸馏实现在线特征学习,性能超越基线
- 适用于标注稀缺场景,适合智能安防系统部署
视频监控系统是智慧城市建设中保障公共安全与管理的关键。作为其基础任务之一,文本到图像的行人检索旨在从图像库中找出最匹配给定文本描述的目标行人。现有方法多依赖目标域充足的标注数据进行监督训练,但在实际应用中,目标域常仅有未标注数据,因标注成本高而难以获取,限制了模型泛化能力。为解决此问题,本文提出一种新型无监督域适应方法——基于图的跨域知识蒸馏(GCKD),用于在跨数据集场景下学习图文行人检索的跨模态特征表示。该方法包含两个核心组件:首先,设计基于图的多模态传播模块,建立视觉与文本样本间的跨域关联;其次,提出对比动量知识蒸馏模块,通过在线知识蒸馏策略学习跨模态特征表示。联合优化两模块后,所提方法在三个公开可用的图文行人检索数据集上均取得优异性能,持续优于当前最优基线。
原文摘要 · Abstract (English)
Video surveillance systems are crucial components for ensuring public safety and management in smart city. As a fundamental task in video surveillance, text-to-image person retrieval aims to retrieve the target person from an image gallery that best matches the given text description. Most existing text-to-image person retrieval methods are trained in a supervised manner that requires sufficient labeled data in the target domain. However, it is common in practice that only unlabeled data is available in the target domain due to the difficulty and cost of data annotation, which limits the generalization of existing methods in practical application scenarios. To address this issue, we propose a novel unsupervised domain adaptation method, termed Graph-Based Cross-Domain Knowledge Distillation (GCKD), to learn the cross-modal feature representation for text-to-image person retrieval in a cross-dataset scenario. The proposed GCKD method consists of two main components. Firstly, a graph-based multi-modal propagation module is designed to bridge the cross-domain correlation among the visual and textual samples. Secondly, a contrastive momentum knowledge distillation module is proposed to learn the cross-modal feature representation using the online knowledge distillation strategy. By jointly optimizing the two modules, the proposed method is able to achieve efficient performance for cross-dataset text-to-image person retrieval. acExtensive experiments on three publicly available text-to-image person retrieval datasets demonstrate the effectiveness of the proposed GCKD method, which consistently outperforms the state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。