arXiv:2601.08807cs.CVcs.AI2026-01中稿 · the 2026 IEEE/CVF …被引 1

用视频超分提升行人重识别的轨迹质量,尤其在跨视角场景下表现更好。

S3-CLIP: Video Super Resolution for Person-ReID

论文配图:S3-CLIP: Video Super Resolution for Person-ReID
图 1 · 摘自论文原文
  • 结合超分网络与任务导向的超分流程,优化视频行人重识别的轨迹质量。
  • 在空地跨视角下达到37.52% mAP,地到空场景下排名指标提升超10%。
  • 首个系统性探索视频超分用于改善行人重识别质量的工作,适合实际部署场景。

大多数行人重识别(ReID)方法将轨迹质量视为次要问题,多数研究集中于基础模型的结构改进。然而,这种忽视在真实复杂场景中带来挑战。本文提出S3-CLIP,一种基于视频超分辨率的CLIP-ReID框架,专为WACV 2026的VReID-XFD挑战设计。该方法融合最新的超分网络与任务驱动的超分流程,适配视频行人重识别任务。据我们所知,这是首个系统性研究视频超分辨率提升行人重识别轨迹质量的工作,尤其在跨视角条件下。实验表明,性能媲美基线,在空地和地空场景下分别取得37.52%和29.16%的mAP。在地到空场景中,Rank-1、Rank-5、Rank-10分别提升11.24%、13.48%、17.98%。

原文摘要 · Abstract (English)

Tracklet quality is often treated as an afterthought in most person re-identification (ReID) methods, with the majority of research presenting architectural modifications to foundational models. Such approaches neglect an important limitation, posing challenges when deploying ReID systems in real-world, difficult scenarios. In this paper, we introduce S3-CLIP, a video super-resolution-based CLIP-ReID framework developed for the VReID-XFD challenge at WACV 2026. The proposed method integrates recent advances in super-resolution networks with task-driven super-resolution pipelines, adapting them to the video-based person re-identification setting. To the best of our knowledge, this work represents the first systematic investigation of video super-resolution as a means of enhancing tracklet quality for person ReID, particularly under challenging cross-view conditions. Experimental results demonstrate performance competitive with the baseline, achieving 37.52% mAP in aerial-to-ground and 29.16% mAP in ground-to-aerial scenarios. In the ground-to-aerial setting, S3-CLIP achieves substantial gains in ranking accuracy, improving Rank-1, Rank-5, and Rank-10 performance by 11.24%, 13.48%, and 17.98%, respectively.

视频超分行人重识别跨视角轨迹增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。