arXiv:2606.02242cs.CVcs.AI2026-06

解决图像与文本行人重识别的优化冲突,实现共享表征统一训练。

Towards Resolving Optimization Conflicts Between Image- and Text-Based Person Re-Identification

论文配图:Towards Resolving Optimization Conflicts Between Image- and Text-Based Person Re-Identification
图 1 · 摘自论文原文
  • 采用分阶段双流训练,用单一视觉编码器支持跨模态检索。
  • 图像重识别预训练显著提升文本重识别泛化能力。
  • 文本监督能同时提升图像和文本任务性能,适合多模态系统开发者。

图像与文本行人重识别(ReID)联合优化受模态差异和训练目标冲突影响,导致共享表征质量不佳。图像重识别关注同一人不同图像间的身份不变性,而文本重识别依赖于描述个体视觉特征的实例特定文本。本文探讨两类任务的本质差异及其优化机制。由于两者常独立研究,一种检索设置的损失函数可能损害另一任务所需的表征质量。为此,我们提出基于单个视觉编码器的解耦式两阶段训练流程,支持图像与文本双重检索,并避免训练过程中的任务干扰。在多种配置下进行充分实验,结果表明:图像重识别预训练可提升对文本数据的泛化能力;在视觉编码器训练阶段引入文本监督,能同时增强图像与文本重识别性能。本工作为统一的行人重识别系统及跨模态检索提供了重要启示。

原文摘要 · Abstract (English)

The joint optimization of image-based (I2I) and text-based (T2I) person re-identification (ReID) is hindered by modality discrepancies and conflicting training objectives, leading to suboptimal shared representations. While I2I ReID focuses on identity-level invariance across images of the same person, T2I ReID is driven by instance-specific textual descriptions tied to unique visual traits. This paper explores the fundamental difference between two ReID tasks and their optimization processes for effective training. Since I2I and T2I ReID are often studied separately, the loss functions optimized for one retrieval setting may negatively affect the representation quality required by the other. Motivated by these findings, we propose a decoupled two-stage training pipeline for learning a shared representation across image and text modalities. The pipeline is based on a single vision encoder that supports both I2I and T2I retrieval while avoiding cross-task interference during training. We provide extensive experiments across multiple configurations, varying domain mixing procedures, learning strategies, and task objectives. We observed that I2I ReID pre-training positively impacts the generalization ability to T2I data. Besides, we find that incorporating textual supervision during the vision encoder training stage enhances both I2I and T2I performance. We believe our insights provide a meaningful step toward unified ReID systems and cross-modal retrieval overall.

行人重识别跨模态共享表征多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。