针对远距离视频行人重识别难题,提出自适应尺度与形状先验框架。
SAS-VPReID: A Scale-Adaptive Framework with Shape Priors for Video-based Person Re-Identification at Extreme Far Distances
- 构建三模块框架:记忆增强视觉主干、多粒度时序建模、先验正则化形状动态
- 在VReID-XFD数据集上排名第一,显著提升极端远距离识别性能
- 适合关注远距离监控与跨摄像头行人追踪的研究者
基于视频的行人重识别(VPReID)旨在从非重叠摄像头拍摄的视频中检索同一行人。在极端远距离场景下,由于分辨率严重退化、视角剧烈变化和不可避免的外观噪声,该任务极具挑战性。为此,本文提出一种融合形状先验的自适应尺度框架SAS-VPReID。该框架由三个互补模块构成:首先,采用记忆增强视觉主干(MEVB),结合CLIP视觉编码器与多代理记忆,提取更具区分性的特征表示;其次,设计多粒度时序建模(MGTM),在多种时间粒度上构建序列并自适应强调跨尺度运动线索;最后,引入先验正则化形状动态(PRSD)捕捉人体结构动态变化。实验表明各模块均有效,最终框架在VReID-XFD基准测试中位列第一。代码已开源于https://github.com/YangQiWei3/SAS-VPReID。
原文摘要 · Abstract (English)
Video-based Person Re-IDentification (VPReID) aims to retrieve the same person from videos captured by non-overlapping cameras. At extreme far distances, VPReID is highly challenging due to severe resolution degradation, drastic viewpoint variation and inevitable appearance noise. To address these issues, we propose a Scale-Adaptive framework with Shape Priors for VPReID, named SAS-VPReID. The framework is built upon three complementary modules. First, we deploy a Memory-Enhanced Visual Backbone (MEVB) to extract discriminative feature representations, which leverages the CLIP vision encoder and multi-proxy memory. Second, we propose a Multi-Granularity Temporal Modeling (MGTM) to construct sequences at multiple temporal granularities and adaptively emphasize motion cues across scales. Third, we incorporate Prior-Regularized Shape Dynamics (PRSD) to capture body structure dynamics. With these modules, our framework can obtain more discriminative feature representations. Experiments on the VReID-XFD benchmark demonstrate the effectiveness of each module and our final framework ranks the first on the VReID-XFD challenge leaderboard. The source code is available at https://github.com/YangQiWei3/SAS-VPReID.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。