提出新方法提升无人机与地面摄像头间行人重识别效果
View-Aware Semantic Alignment for Aerial-Ground Person Re-Identification

- 设计视角感知模块,生成能捕捉视角特性的语义查询
- 在CARGO数据集上提升10.06% mAP,显著优于现有方法
- 适合关注跨视角行人识别的科研与工程人员
由于无人机与固定摄像头之间视角差异巨大,航拍-地面行人重识别(AGPReID)仍极具挑战。现有方法多采用视图不变范式,对齐跨视角共享特征以增强鲁棒性,但视图不变性强制进行局部对齐,忽略视角特异性线索和判别性身份信息。为此,本文提出视图感知语义对齐框架ViSA,包含专家驱动的令牌生成模块(ETGM)和双分支局部融合模块(DLFM)。前者构建一组视角感知专家,生成自适应的语义查询以感知视角特异性模式;后者利用图推理提取并对齐对不同专家响应的局部区域。在三个基准数据集AG-ReID.v2、CARGO和LAGPeR上的大量实验表明,ViSA持续取得更优性能,在具有挑战性的CARGO跨视角协议上实现10.06% mAP提升。代码已公开于https://github.com/Cat-Zero/ViSA。
原文摘要 · Abstract (English)
Aerial-Ground Person Re-Identification (AGPReID) remains highly challenging due to drastic viewpoint variations between drones and fixed cameras. Existing methods typically follow a view-invariant paradigm, aligning shared features across views to achieve robustness. However, view-invariant inherently enforces part-level alignment, which ignores view-specific cues and discriminative identity information. To this end, this work proposes ViSA (View-aware Semantic Alignment), a view-aware framework that achieves cross-view semantic consistency containing an Expert-driven Token Generation Module (ETGM) and a Dual-branch Local Fusion Module (DLFM). Technically, the former constructs a set of view-aware experts to generate adaptive semantic queries that perceive viewpoint-specific patterns, while the latter leverages graph reasoning to extract and align local regions responsive to different experts. Extensive experiments on three AGPReID benchmarks including AG-ReID.v2, CARGO and LAGPeR demonstrate that ViSA consistently achieves superior performance, with a notable 10.06\% mAP improvement on the challenging CARGO cross-view protocol. The code is available at \href{https://github.com/Cat-Zero/ViSA}{https://github.com/Cat-Zero/ViSA}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。