arXiv:2602.05785cs.CVcs.AI2026-02被引 1

用文本增强单摄像头数据,提升跨域行人重识别泛化能力

ReText: Text Boosts Generalization in Image-Based Person Re-identification

  • 混合多相机与单相机数据,通过文本补充语义信息
  • 三任务联合优化:重识别、图文匹配、文本引导重建
  • 在多个跨域基准上超越现有方法,适合实际部署场景

通用图像行人重识别旨在无需重新训练的情况下,跨摄像头识别未见域中的个体。尽管已有方法通过复杂架构缓解领域差异,但近期研究发现,风格多样化的单摄像头数据更利于提升泛化性能。此类数据易获取,但缺乏视角变化带来的复杂性。我们提出ReText,一种在多相机与单相机数据混合基础上训练的新方法,其中单相机数据通过文本描述增强语义信息。训练时,ReText联合优化三项任务:(1) 多相机上的重识别,(2) 图文匹配,(3) 单相机数据上由文本引导的图像重建。实验表明,ReText在跨域重识别基准上实现强泛化性能,显著优于现有先进方法。据我们所知,这是首个探索在图像行人重识别中对多相机与单相机数据混合进行多模态联合学习的工作。

原文摘要 · Abstract (English)

Generalizable image-based person re-identification (Re-ID) aims to recognize individuals across cameras in unseen domains without retraining. While multiple existing approaches address the domain gap through complex architectures, recent findings indicate that better generalization can be achieved by stylistically diverse single-camera data. Although this data is easy to collect, it lacks complexity due to minimal cross-view variation. We propose ReText, a novel method trained on a mixture of multi-camera Re-ID data and single-camera data, where the latter is complemented by textual descriptions to enrich semantic cues. During training, ReText jointly optimizes three tasks: (1) Re-ID on multi-camera data, (2) image-text matching, and (3) image reconstruction guided by text on single-camera data. Experiments demonstrate that ReText achieves strong generalization and significantly outperforms state-of-the-art methods on cross-domain Re-ID benchmarks. To the best of our knowledge, this is the first work to explore multimodal joint learning on a mixture of multi-camera and single-camera data in image-based person Re-ID.

行人重识别多模态学习泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。