用视觉变压器识别远距离、高空下人体轮廓,持久稳定。
Unconstrained Body Recognition at Altitude and Range: Comparing Four Approaches
- 基于ViT和Swin-ViT构建人体轮廓识别模型,学习长期稳定特征。
- 在超190万张图像上训练,跨9个数据集验证,最远识别达1000米。
- 适合无人机监控、长时序行人追踪等真实场景应用。
本研究探讨四种不同方法在长时序人体识别中的表现,聚焦于随时间稳定的体态特征,而非临时特征(如服装)。提出基于视觉变压器(ViT)的体态识别模型(BIDDS)与Swin-ViT模型(Swin-BIDDS),并改进了先前基于核心残差网络的LCRIM与NLCRIM方法。所有模型均在包含约5000个身份、超过190万张图像的大型多样化数据集上训练,覆盖9个数据库。性能在标准重识别基准数据集(MARS、MSMT17、Outdoor Gait、DeepChange)及一个非约束数据集上评估,该数据集包含从近距到1000米距离、无人机高空拍摄及服装变化的图像。对比分析揭示了不同主干架构与输入图像尺寸对实际复杂环境下体态识别性能的影响。
原文摘要 · Abstract (English)
This study presents an investigation of four distinct approaches to long-term person identification using body shape. Unlike short-term re-identification systems that rely on temporary features (e.g., clothing), we focus on learning persistent body shape characteristics that remain stable over time. We introduce a body identification model based on a Vision Transformer (ViT) (Body Identification from Diverse Datasets, BIDDS) and on a Swin-ViT model (Swin-BIDDS). We also expand on previous approaches based on the Linguistic and Non-linguistic Core ResNet Identity Models (LCRIM and NLCRIM), but with improved training. All models are trained on a large and diverse dataset of over 1.9 million images of approximately 5k identities across 9 databases. Performance was evaluated on standard re-identification benchmark datasets (MARS, MSMT17, Outdoor Gait, DeepChange) and on an unconstrained dataset that includes images at a distance (from close-range to 1000m), at altitude (from an unmanned aerial vehicle, UAV), and with clothing change. A comparative analysis across these models provides insights into how different backbone architectures and input image sizes impact long-term body identification performance across real-world conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。