解决可见光与红外跨模态持续识别中的知识干扰问题。
CKDA: Cross-modality Knowledge Disentanglement and Alignment for Visible-Infrared Lifelong Person Re-identification
- 提出分离特定模态与共用知识的双提示模块,避免知识互扰。
- 在四个数据集上实现优于现有方法的平均性能,提升显著。
- 适合长期跨昼夜监控、多模态行人识别场景应用。
持续行人重识别(LReID)旨在利用不同场景下连续采集的个体数据匹配同一人。为实现全天候跨昼夜匹配,可见光-红外持续行人重识别(VI-LReID)需在可见光与红外模态数据上顺序训练,并追求所有数据上的平均性能。现有方法通常采用跨模态知识蒸馏缓解旧知识的灾难性遗忘,但忽略了特定模态知识获取与共用知识抗遗忘之间的相互干扰,导致协同遗忘。为此,本文提出跨模态知识解耦与对齐方法(CKDA),显式平衡地分离并保留特定模态与共用知识。具体而言,设计了模态共用提示(MCP)与模态特定提示(MSP)模块,以分离并净化各模态特有的判别信息,避免知识间干扰;同时引入跨模态知识对齐(CKA)模块,基于双模态原型,在独立的模态内外特征空间中平衡对齐新旧知识。在四个基准数据集上的大量实验验证了所提方法的有效性与优越性。
原文摘要 · Abstract (English)
Lifelong person Re-IDentification (LReID) aims to match the same person employing continuously collected individual data from different scenarios. To achieve continuous all-day person matching across day and night, Visible-Infrared Lifelong person Re-IDentification (VI-LReID) focuses on sequential training on data from visible and infrared modalities and pursues average performance over all data. To this end, existing methods typically exploit cross-modal knowledge distillation to alleviate the catastrophic forgetting of old knowledge. However, these methods ignore the mutual interference of modality-specific knowledge acquisition and modality-common knowledge anti-forgetting, where conflicting knowledge leads to collaborative forgetting. To address the above problems, this paper proposes a Cross-modality Knowledge Disentanglement and Alignment method, called CKDA, which explicitly separates and preserves modality-specific knowledge and modality-common knowledge in a balanced way. Specifically, a Modality-Common Prompting (MCP) module and a Modality-Specific Prompting (MSP) module are proposed to explicitly disentangle and purify discriminative information that coexists and is specific to different modalities, avoiding the mutual interference between both knowledge. In addition, a Cross-modal Knowledge Alignment (CKA) module is designed to further align the disentangled new knowledge with the old one in two mutually independent inter- and intra-modality feature spaces based on dual-modality prototypes in a balanced manner. Extensive experiments on four benchmark datasets verify the effectiveness and superiority of our CKDA against state-of-the-art methods. The source code of this paper is available at https://github.com/PKU-ICST-MIPL/CKDA-AAAI2026.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。