arXiv:2604.04183cs.CV2026-04

用大模型提升远距离视频行人重识别准确率

Scale-Aware Vision-Language Adaptation for Extreme Far-Distance Video Person Re-identification

  • 升级视觉主干并采用选择性微调,增强大模型稳定性
  • 在极端远距离下达到35.73的总体mAP,显著优于基线
  • 适合研究长距离监控、无人机行人识别的学者使用

极端远距离视频行人重识别因尺度压缩、分辨率下降、运动模糊和空地视角差异而极具挑战。随着摄像头高度与目标距离增加,基于近距离图像训练的模型性能大幅退化。本文研究如何将大规模视觉-语言模型适配于此类场景。从基于CLIP的基线出发,将视觉主干由ViT-B/16升级为ViT-L/14,并引入骨干感知的选择性微调以稳定大模型适应过程。针对噪声多、分辨率低的轨迹片段,设计轻量级时序注意力池化机制,抑制劣化帧并强化关键观测。保留适配器与提示条件跨视角学习以缓解空地域偏移,并通过改进优化与k-互近邻重排序进一步提升检索效果。在DetReIDX压力测试基准上,本方法取得A2G 46.69、G2A 41.23、A2A 22.98的mAP,总体mAP达35.73。结果表明,结合稳定性适配的大规模视觉-语言主干可显著提升极端远距离视频行人重识别的鲁棒性。

原文摘要 · Abstract (English)

Extreme far-distance video person re-identification (ReID) is particularly challenging due to scale compression, resolution degradation, motion blur, and aerial-ground viewpoint mismatch. As camera altitude and subject distance increase, models trained on close-range imagery degrade significantly. In this work, we investigate how large-scale vision-language models can be adapted to operate reliably under these conditions. Starting from a CLIP-based baseline, we upgrade the visual backbone from ViT-B/16 to ViT-L/14 and introduce backbone-aware selective fine-tuning to stabilize adaptation of the larger transformer. To address noisy and low-resolution tracklets, we incorporate a lightweight temporal attention pooling mechanism that suppresses degraded frames and emphasizes informative observations. We retain adapter-based and prompt-conditioned cross-view learning to mitigate aerial-ground domain shifts, and further refine retrieval using improved optimization and k-reciprocal re-ranking. Experiments on the DetReIDX stress-test benchmark show that our approach achieves mAP scores of 46.69 (A2G), 41.23 (G2A), and 22.98 (A2A), corresponding to an overall mAP of 35.73. These results show that large-scale vision-language backbones, when combined with stability-focused adaptation, significantly enhance robustness in extreme far-distance video person ReID.

视频重识别视觉语言模型远距离识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。