arXiv:2604.19218cs.CV2026-04

让机器像人一样思考再匹配,提升跨场景身份识别能力

Thinking Before Matching: A Reinforcement Reasoning Paradigm Towards General Person Re-Identification

论文配图:Thinking Before Matching: A Reinforcement Reasoning Paradigm Towards General Person Re-Identification
图 1 · 摘自论文原文
  • 用思维链方式训练模型理解身份特征,不依赖大量标注数据
  • 仅用14.3K条非平凡数据就达到主流方法水平,数据量减少79.1%
  • 能自动生成推理解释,适合需要可解释性的实际应用

在行人重识别(ReID)中,学习具备跨场景泛化能力的身份判别表征已成为关键目标。然而主流感知驱动范式倾向于从海量标注数据中拟合匹配,而非理解身份因果线索,导致表征对多种干扰脆弱。本文提出一种新型推理驱动范式ReID-R,通过将思维链(CoT)融入ReID流程,实现显式的身份理解与推理。具体包含两个阶段:(i) 判别性推理预热,模型在无标签思维链下训练,获得身份感知的特征理解;(ii) 高效强化学习,设计非平凡采样策略构建具有场景泛化能力的数据。在此基础上,利用高质量奖励信号引导模型聚焦于身份相关线索,实现精准推理与正确响应。多个ReID基准上的大量实验表明,ReID-R仅需14.3K条非平凡数据(现有数据规模的20.9%),即可达到与先进方法相当的身份判别性能。此外,得益于内在推理机制,ReID-R还能为结果提供高质量解释。

原文摘要 · Abstract (English)

Learning identity-discriminative representations with multi-scene generality has become a critical objective in person re-identification (ReID). However, mainstream perception-driven paradigms tend to identify fitting from massive annotated data rather than identity-causal cues understanding, which presents a fragile representation against multiple disruptions. In this work, ReID-R is proposed as a novel reasoning-driven paradigm that achieves explicit identity understanding and reasoning by incorporating chain-of-thought into the ReID pipeline. Specifically, ReID-R consists of a two-stage contribution: (i) Discriminative reasoning warm-up, where a model is trained in a CoT label-free manner to acquire identity-aware feature understanding; and (ii) Efficient reinforcement learning, which proposes a non-trivial sampling to construct scene-generalizable data. On this basis, ReID-R leverages high-quality reward signals to guide the model toward focusing on ID-related cues, achieving accurate reasoning and correct responses. Extensive experiments on multiple ReID benchmarks demonstrate that ReID-R achieves competitive identity discrimination as superior methods using only 14.3K non-trivial data (20.9% of the existing data scale). Furthermore, benefit from inherent reasoning, ReID-R can provide high-quality interpretation for results.

身份识别强化学习可解释性思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。