arXiv:2505.17821cs.CV2025-05中稿 · IEEE Transactions …被引 12

用文本提示学习统一多光谱特征,提升跨谱识别准确率。

ICPL-ReID: Identity-Conditional Prompt Learning for Multi-Spectral Object Re-Identification

  • 通过可学习文本提示建模身份语义中心,在线对齐不同光谱特征。
  • 在无明确描述时,用身份原型约束提示学习,提升稳定性。
  • 适合小样本多光谱数据,尤其适用于智慧城市与交通场景。

多光谱目标再识别为智慧城市场景和智能交通提供了新视角,有效应对复杂光照与恶劣天气挑战。然而,异质光谱间的差异使互补与矛盾信息难以高效利用。现有方法多依赖复杂的模态交互模块,缺乏对光谱信息的细粒度语义理解(如文本描述、部位掩码、物体关键点)。为此,我们提出新型身份条件文本提示学习框架(ICPL),利用CLIP强大的跨模态对齐能力,统一不同光谱的视觉特征。首先,提出在线提示学习,使用可学习文本提示作为身份级语义中心,在线桥接不同光谱的身份语义。其次,在缺乏具体文本描述时,设计多光谱身份条件模块,以身份原型作为光谱身份条件,约束提示学习。同时,构建对齐环路,互优化可学习文本提示与光谱视觉编码器,避免在线提示学习破坏预训练的图文对齐分布。此外,为适应小规模多光谱数据并缓解光谱间风格差异,提出多光谱适配器,采用低秩适配方法学习光谱特异性特征。在5个基准测试(RGBNT201、Market-MM、MSVR310、RGBN300、RGBNT100)上的全面实验表明,该方法优于当前最优方法。

原文摘要 · Abstract (English)

Multi-spectral object re-identification (ReID) brings a new perception perspective for smart city and intelligent transportation applications, effectively addressing challenges from complex illumination and adverse weather. However, complex modal differences between heterogeneous spectra pose challenges to efficiently utilizing complementary and discrepancy of spectra information. Most existing methods fuse spectral data through intricate modal interaction modules, lacking fine-grained semantic understanding of spectral information (\textit{e.g.}, text descriptions, part masks, and object keypoints). To solve this challenge, we propose a novel Identity-Conditional text Prompt Learning framework (ICPL), which exploits the powerful cross-modal alignment capability of CLIP, to unify different spectral visual features from text semantics. Specifically, we first propose the online prompt learning using learnable text prompt as the identity-level semantic center to bridge the identity semantics of different spectra in online manner. Then, in lack of concrete text descriptions, we propose the multi-spectral identity-condition module to use identity prototype as spectral identity condition to constraint prompt learning. Meanwhile, we construct the alignment loop mutually optimizing the learnable text prompt and spectral visual encoder to avoid online prompt learning disrupting the pre-trained text-image alignment distribution. In addition, to adapt to small-scale multi-spectral data and mitigate style differences between spectra, we propose multi-spectral adapter that employs a low-rank adaption method to learn spectra-specific features. Comprehensive experiments on 5 benchmarks, including RGBNT201, Market-MM, MSVR310, RGBN300, and RGBNT100, demonstrate that the proposed method outperforms the state-of-the-art methods.

多光谱再识别提示学习CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。