arXiv:2510.22268cs.CV2025-10NeurIPS被引 5

解决无人机与地面视角下行人重识别的几何与语义错位问题

GSAlign: Geometric and Semantic Alignment Network for Aerial-Ground Person Re-Identification

  • 用可学习的薄板样条模块自适应校正极端视角带来的几何畸变
  • 通过动态对齐模块生成可见区域掩码,缓解遮挡和部分观测影响
  • 在CARGO数据集上提升mAP 18.8%、Rank-1 16.8%,适合跨视角匹配场景

航拍-地面行人重识别(AG-ReID)是一项新兴且极具挑战性的任务,旨在匹配由无人机(UAV)和地面监控摄像头拍摄的视角差异极大的行人图像。由于视角差异巨大、遮挡严重以及航拍与地面图像间的域差距,该任务面临严峻挑战。尽管已有方法在跨视角表征学习方面取得进展,但在处理严重姿态变化和空间错位方面仍存在局限。为此,我们提出专为AG-ReID设计的几何与语义对齐网络(GSAlign)。GSAlign引入两个关键组件:可学习薄板样条(LTPS)模块和动态对齐模块(DAM),联合解决航拍-地面匹配中的几何失真与语义错位问题。LTPS模块基于一组学习的关键点自适应地扭曲行人特征,有效补偿因极端视角变化引起的几何变形;同时,DAM估计感知可见性的表示掩码,在语义层面突出可见身体区域,从而缓解遮挡和部分观测对跨视图对应关系的负面影响。在CARGO数据集上,采用四种匹配协议的全面评估表明,GSAlign显著优于先前最先进方法,航拍-地面设置下mAP提升+18.8%,Rank-1准确率提升+16.8%。

原文摘要 · Abstract (English)

Aerial-Ground person re-identification (AG-ReID) is an emerging yet challenging task that aims to match pedestrian images captured from drastically different viewpoints, typically from unmanned aerial vehicles (UAVs) and ground-based surveillance cameras. The task poses significant challenges due to extreme viewpoint discrepancies, occlusions, and domain gaps between aerial and ground imagery. While prior works have made progress by learning cross-view representations, they remain limited in handling severe pose variations and spatial misalignment. To address these issues, we propose a Geometric and Semantic Alignment Network (GSAlign) tailored for AG-ReID. GSAlign introduces two key components to jointly tackle geometric distortion and semantic misalignment in aerial-ground matching: a Learnable Thin Plate Spline (LTPS) Module and a Dynamic Alignment Module (DAM). The LTPS module adaptively warps pedestrian features based on a set of learned keypoints, effectively compensating for geometric variations caused by extreme viewpoint changes. In parallel, the DAM estimates visibility-aware representation masks that highlight visible body regions at the semantic level, thereby alleviating the negative impact of occlusions and partial observations in cross-view correspondence. A comprehensive evaluation on CARGO with four matching protocols demonstrates the effectiveness of GSAlign, achieving significant improvements of +18.8\% in mAP and +16.8\% in Rank-1 accuracy over previous state-of-the-art methods on the aerial-ground setting.

行人重识别跨视角几何对齐视觉匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。