融合局部细节与全局语义,提升行人重识别精度
DRFormer: A Dual-Regularized Bidirectional Transformer for Person Re-identification
- 设计双向双正则化Transformer,协同视觉基础模型与视觉语言模型
- 在5个基准上达到领先性能,有效缓解遮挡与姿态变化影响
- 适合关注多模态特征融合的行人识别研究者
细粒度判别性特征与全局语义信息均有助于解决行人重识别中的遮挡和姿态变化问题。视觉基础模型(如DINO)擅长挖掘局部纹理,而视觉语言模型(如CLIP)能捕捉强全局语义差异。现有方法多依赖单一范式,忽视两者整合潜力。本文分析两类模型的互补作用,提出双正则化双向Transformer(DRFormer),通过双重正则化机制确保特征多样性并平衡两者贡献。在五个基准上的大量实验表明,该方法有效融合局部与全局表征,性能媲美当前最优方法。
原文摘要 · Abstract (English)
Both fine-grained discriminative details and global semantic features can contribute to solving person re-identification challenges, such as occlusion and pose variations. Vision foundation models (\textit{e.g.}, DINO) excel at mining local textures, and vision-language models (\textit{e.g.}, CLIP) capture strong global semantic difference. Existing methods predominantly rely on a single paradigm, neglecting the potential benefits of their integration. In this paper, we analyze the complementary roles of these two architectures and propose a framework to synergize their strengths by a \textbf{D}ual-\textbf{R}egularized Bidirectional \textbf{Transformer} (\textbf{DRFormer}). The dual-regularization mechanism ensures diverse feature extraction and achieves a better balance in the contributions of the two models. Extensive experiments on five benchmarks show that our method effectively harmonizes local and global representations, achieving competitive performance against state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。