arXiv:2509.22331cs.CVcs.AI2025-09被引 2

构建跨模态超图,提升行人属性识别准确率

Pedestrian Attribute Recognition via Hierarchical Cross-Modality HyperGraph Learning

  • 通过多模态知识图谱建模属性与视觉特征关系
  • 在多个基准数据集上显著提升识别性能
  • 适合关注行人识别与知识融合的研究者

当前行人属性识别(PAR)算法通常聚焦于将视觉特征映射到语义标签,或通过融合视觉与属性信息来增强学习。然而,这些方法未能充分挖掘属性知识和上下文信息以实现更精准的识别。尽管近期研究开始引入属性文本作为额外输入以增强视觉与语义之间的关联,但此类方法仍处于初级阶段。为此,本文提出构建一个多模态知识图谱,用于挖掘局部视觉特征与文本之间的关系,以及属性与广泛视觉上下文样本之间的关系。具体而言,提出一种有效的多模态知识图谱构建方法,全面考虑属性间的相互关系及其与视觉标记的关系。为有效建模这些关系,本文引入一种知识图谱引导的跨模态超图学习框架,以增强标准行人属性识别框架。在多个PAR基准数据集上的综合实验充分证明了所提知识图谱在PAR任务中的有效性,为知识引导的行人属性识别奠定了坚实基础。论文源码将发布于 https://github.com/Event-AHU/OpenPAR

原文摘要 · Abstract (English)

Current Pedestrian Attribute Recognition (PAR) algorithms typically focus on mapping visual features to semantic labels or attempt to enhance learning by fusing visual and attribute information. However, these methods fail to fully exploit attribute knowledge and contextual information for more accurate recognition. Although recent works have started to consider using attribute text as additional input to enhance the association between visual and semantic information, these methods are still in their infancy. To address the above challenges, this paper proposes the construction of a multi-modal knowledge graph, which is utilized to mine the relationships between local visual features and text, as well as the relationships between attributes and extensive visual context samples. Specifically, we propose an effective multi-modal knowledge graph construction method that fully considers the relationships among attributes and the relationships between attributes and vision tokens. To effectively model these relationships, this paper introduces a knowledge graph-guided cross-modal hypergraph learning framework to enhance the standard pedestrian attribute recognition framework. Comprehensive experiments on multiple PAR benchmark datasets have thoroughly demonstrated the effectiveness of our proposed knowledge graph for the PAR task, establishing a strong foundation for knowledge-guided pedestrian attribute recognition. The source code of this paper will be released on https://github.com/Event-AHU/OpenPAR

行人属性识别多模态学习知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。