arXiv:2604.15090cs.CV2026-04

用语义文本增强行人重识别,抗光照和换装干扰

Beyond Visual Cues: Semantic-Driven Token Filtering and Expert Routing for Anytime Person ReID

论文配图:Beyond Visual Cues: Semantic-Driven Token Filtering and Expert Routing for Anytime Person ReID
图 1 · 摘自论文原文
  • 用大视觉语言模型生成身份不变的语义文本,指导视觉特征筛选
  • 在AT-USTC数据集上超越现有方法,跨5个基准测试表现优异
  • 适合需要应对复杂环境变化的行人识别场景

任意时间行人重识别(AT-ReID)需在任意条件下可靠检索目标个体,涵盖光照变化(白天/夜晚)及长时间跨度的着装改变。现有方法严重依赖纯视觉特征,易受环境与时间因素影响,导致性能显著下降。本文提出语义驱动的令牌过滤与专家路由框架(STFER),利用大视觉语言模型(LVLM)生成具有身份一致性的语义文本,提供对服装变化和跨模态(RGB与红外)切换均鲁棒的身份判别特征。具体地,通过指令引导LVLM生成捕捉生物恒定特征的内在语义文本,该文本用于语义驱动的视觉令牌过滤(SVTF),增强有效视觉区域并抑制冗余背景噪声;同时用于语义驱动的专家路由(SER),实现更鲁棒的多场景门控。在任意时间重识别数据集(AT-USTC)上的大量实验表明,本模型达到当前最优性能。此外,基于AT-USTC训练的模型在5个主流重识别基准上评估,展现出卓越的泛化能力,结果极具竞争力。代码即将发布。

原文摘要 · Abstract (English)

Any-Time Person Re-identification (AT-ReID) necessitates the robust retrieval of target individuals under arbitrary conditions, encompassing both modality shifts (daytime and nighttime) and extensive clothing-change scenarios, ranging from short-term to long-term intervals. However, existing methods are highly relying on pure visual features, which are prone to change due to environmental and time factors, resulting in significantly performance deterioration under scenarios involving illumination caused modality shifts or cloth-change. In this paper, we propose Semantic-driven Token Filtering and Expert Routing (STFER), a novel framework that leverages the ability of Large Vision-Language Models (LVLMs) to generate identity consistency text, which provides identity-discriminative features that are robust to both clothing variations and cross-modality shifts between RGB and IR. Specifically, we employ instructions to guide the LVLM in generating identity-intrinsic semantic text that captures biometric constants for the semantic model driven. The text token is further used for Semantic-driven Visual Token Filtering (SVTF), which enhances informative visual regions and suppresses redundant background noise. Meanwhile, the text token is also used for Semantic-driven Expert Routing (SER), which integrates the semantic text into expert routing, resulting in more robust multi-scenario gating. Extensive experiments on the Any-Time ReID dataset (AT-USTC) demonstrate that our model achieves state-of-the-art results. Moreover, the model trained on AT-USTC was evaluated across 5 widely-used ReID benchmarks demonstrating superior generalization capabilities with highly competitive results. Our code will be available soon.

行人重识别语义引导多模态视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。