arXiv:2509.03032cs.CV2025-09被引 2

让背景也说话:用语言增强对抗学习提升行人重识别

Background Matters Too: A Language-Enhanced Adversarial Framework for Person Re-Identification

  • 双分支跨模态架构,同时建模前景与背景语义
  • 通过语义对齐与对抗学习,显著抑制背景干扰
  • 无需人工标注,适合复杂遮挡场景下的身份识别

行人重识别面临两大挑战:准确定位前景目标并抑制背景噪声,以及从目标区域提取细粒度特征。现有视觉方法依赖人工标注且难以应对复杂遮挡;近期多模态方法虽引入语义线索,但仅关注前景而忽略背景信息。受人类感知启发,本文认为背景语义与前景同样重要——人会主动忽略背景干扰以聚焦目标外观。为此,提出端到端框架,在双分支跨模态特征提取中联合建模前景与背景。设计内部语义对齐与跨语义对抗学习策略:对齐同义的跨模态特征,同时惩罚前景与背景特征间的相似性,增强模型判别力。该策略促使网络主动抑制噪声背景,强化对身份相关前景的关注。在两个完整和两个遮挡型ReID基准上的实验表明,本方法效果优于或匹配当前最先进水平。

原文摘要 · Abstract (English)

Person re-identification faces two core challenges: precisely locating the foreground target while suppressing background noise and extracting fine-grained features from the target region. Numerous visual-only approaches address these issues by partitioning an image and applying attention modules, yet they rely on costly manual annotations and struggle with complex occlusions. Recent multimodal methods, motivated by CLIP, introduce semantic cues to guide visual understanding. However, they focus solely on foreground information, but overlook the potential value of background cues. Inspired by human perception, we argue that background semantics are as important as the foreground semantics in ReID, as humans tend to eliminate background distractions while focusing on target appearance. Therefore, this paper proposes an end-to-end framework that jointly models foreground and background information within a dual-branch cross-modal feature extraction pipeline. To help the network distinguish between the two domains, we propose an intra-semantic alignment and inter-semantic adversarial learning strategy. Specifically, we align visual and textual features that share the same semantics across domains, while simultaneously penalizing similarity between foreground and background features to enhance the network's discriminative power. This strategy drives the model to actively suppress noisy background regions and enhance attention toward identity-relevant foreground cues. Comprehensive experiments on two holistic and two occluded ReID benchmarks demonstrate the effectiveness and generality of the proposed method, with results that match or surpass those of current state-of-the-art approaches.

行人重识别多模态学习对抗学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。