纠正高空与地面视角差异导致的相似性扭曲,提升跨视角行人重识别准确率。
Rectifying Geometry-Induced Similarity Distortions for Real-World Aerial-Ground Person Re-Identification
- 设计轻量级模块GIQT,根据相机几何关系修正查询与键的相似性计算
- 在四个基准上实现更优鲁棒性,极端视角下性能显著提升
- 适合关注跨视角重识别、几何畸变校正的研究者
高空-地面行人重识别(AG-ReID)面临空中与地面摄像头间极端视角和距离差异的挑战,导致严重的几何失真,并破坏了跨视角共享相似性空间的假设。现有方法多依赖几何感知特征学习或外观条件提示,但隐含假设注意力机制中几何不变的点积相似性在大视角和尺度变化下依然可靠。本文认为该假设不成立:极端相机几何会系统性扭曲查询-键相似性空间,损害基于注意力的匹配效果,即使特征表示部分对齐。为此,提出几何诱导查询-键变换(GIQT),一个轻量级低秩模块,通过将查询-键交互显式关联相机几何来修正相似性空间。不同于修改特征或注意力形式,GIQT直接调整相似性计算以补偿主导的几何诱导各向异性失真。在此局部相似性校正基础上,进一步引入几何条件提示生成机制,从相机几何直接获取全局、视图自适应的表征先验。在四个高空-地面行人重识别基准上的实验表明,所提框架在极端且未见过的几何条件下均显著提升鲁棒性,同时相比最先进方法引入的计算开销极小。
原文摘要 · Abstract (English)
Aerial-ground person re-identification (AG-ReID) is fundamentally challenged by extreme viewpoint and distance discrepancies between aerial and ground cameras, which induce severe geometric distortions and invalidate the assumption of a shared similarity space across views. Existing methods primarily rely on geometry-aware feature learning or appearance-conditioned prompting, while implicitly assuming that the geometry-invariant dot-product similarity used in attention mechanisms remains reliable under large viewpoint and scale variations. We argue that this assumption does not hold. Extreme camera geometry systematically distorts the query-key similarity space and degrades attention-based matching, even when feature representations are partially aligned. To address this issue, we introduce Geometry-Induced Query-Key Transformation (GIQT), a lightweight low-rank module that explicitly rectifies the similarity space by conditioning query-key interactions on camera geometry. Rather than modifying feature representations or the attention formulation itself, GIQT adapts the similarity computation to compensate for dominant geometry-induced anisotropic distortions. Building on this local similarity rectification, we further incorporate a geometry-conditioned prompt generation mechanism that provides global, view-adaptive representation priors derived directly from camera geometry.Experiments on four aerial-ground person re-identification benchmarks demonstrate that the proposed framework consistently improves robustness under extreme and previously unseen geometric conditions, while introducing minimal computational overhead compared to state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。