arXiv:2505.15825cs.CVcs.AI2025-05被引 19

用张量融合与多线性学习提升行人重识别准确率

Multilinear subspace learning for person re-identification based fusion of high order tensor features

  • 将CNN与LOMO特征融合为统一张量,解决维度不一致问题
  • 在VIPeR、GRID、PRID450S上均超越当前最优方法
  • 适合需要高精度行人匹配的监控系统研发者

视频监控图像分析是计算机视觉中的挑战性领域,其中行人重识别(PRe-ID)尤为困难。该任务旨在通过鲁棒的行人图像描述,在多摄像头网络中识别并追踪已检测到的目标个体。近期研究的成功主要归功于有效的特征提取与表示,以及对这些特征的强大学习能力。为此,本文提出高阶特征融合(HDFF)方法,将卷积神经网络(CNN)和局部最大出现(LOMO)两种强大特征建模为多维数据,并引入新的张量融合方案,将不同维度的特征整合至单一张量中。为进一步提升准确性,采用张量跨视图二次分析(TXQDA)进行多线性子空间学习,随后使用余弦相似度进行匹配。TXQDA有效降低了高阶张量数据的高维性,同时保持学习性能。在三个常用数据集VIPeR、GRID和PRID450S上的实验验证了本方法的有效性,结果表明其优于近期最先进方法。

原文摘要 · Abstract (English)

Video surveillance image analysis and processing is a challenging field in computer vision, with one of its most difficult tasks being Person Re-Identification (PRe-ID). PRe-ID aims to identify and track target individuals who have already been detected in a network of cameras, using a robust description of their pedestrian images. The success of recent research in person PRe-ID is largely due to effective feature extraction and representation, as well as the powerful learning of these features to reliably discriminate between pedestrian images. To this end, two powerful features, Convolutional Neural Networks (CNN) and Local Maximal Occurrence (LOMO), are modeled on multidimensional data using the proposed method, High-Dimensional Feature Fusion (HDFF). Specifically, a new tensor fusion scheme is introduced to leverage and combine these two types of features in a single tensor, even though their dimensions are not identical. To enhance the system's accuracy, we employ Tensor Cross-View Quadratic Analysis (TXQDA) for multilinear subspace learning, followed by cosine similarity for matching. TXQDA efficiently facilitates learning while reducing the high dimensionality inherent in high-order tensor data. The effectiveness of our approach is verified through experiments on three widely-used PRe-ID datasets: VIPeR, GRID, and PRID450S. Extensive experiments demonstrate that our approach outperforms recent state-of-the-art methods.

行人重识别张量融合多线性学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。