首个融合RGB与事件相机的行人属性识别数据集与模型,提升光照和运动下的识别性能。
RGB-Event based Pedestrian Attribute Recognition: A Benchmark Dataset and An Asymmetric RWKV Fusion Framework
- 提出基于事件相机的异构融合框架,利用事件流增强低光/高速场景表征
- 构建10万对样本的EventPAR数据集,覆盖50个外观属性与6种情绪维度
- 适用于智能监控、自动驾驶等需要鲁棒感知的多模态场景
现有行人属性识别方法主要依赖RGB摄像头,受限于光照敏感和运动模糊等问题。同时,当前研究多关注外貌特征,缺乏对情绪维度的探索。本文提出一种新型多模态RGB-事件相机行人属性识别任务,借鉴事件相机在低光、高速、低功耗条件下的优势。我们构建首个大规模多模态数据集EventPAR,包含10万对配对的RGB-事件样本,涵盖50个外观属性及6种人类情绪,覆盖多样场景与季节变化。通过在该数据集上重新训练并评估主流PAR模型,建立全面基准,为后续研究提供数据与算法基础。此外,提出基于RWKV的多模态行人属性识别框架,包含RWKV视觉编码器与非对称融合模块。在所提数据集及两个模拟数据集(MARS-Attribute、DukeMTMC-VID-Attribute)上的实验均取得领先性能。源代码与数据集将公开于https://github.com/Event-AHU/OpenPAR。
原文摘要 · Abstract (English)
Existing pedestrian attribute recognition methods are generally developed based on RGB frame cameras. However, these approaches are constrained by the limitations of RGB cameras, such as sensitivity to lighting conditions and motion blur, which hinder their performance. Furthermore, current attribute recognition primarily focuses on analyzing pedestrians' external appearance and clothing, lacking an exploration of emotional dimensions. In this paper, we revisit these issues and propose a novel multi-modal RGB-Event attribute recognition task by drawing inspiration from the advantages of event cameras in low-light, high-speed, and low-power consumption. Specifically, we introduce the first large-scale multi-modal pedestrian attribute recognition dataset, termed EventPAR, comprising 100K paired RGB-Event samples that cover 50 attributes related to both appearance and six human emotions, diverse scenes, and various seasons. By retraining and evaluating mainstream PAR models on this dataset, we establish a comprehensive benchmark and provide a solid foundation for future research in terms of data and algorithmic baselines. In addition, we propose a novel RWKV-based multi-modal pedestrian attribute recognition framework, featuring an RWKV visual encoder and an asymmetric RWKV fusion module. Extensive experiments are conducted on our proposed dataset as well as two simulated datasets (MARS-Attribute and DukeMTMC-VID-Attribute), achieving state-of-the-art results. The source code and dataset will be released on https://github.com/Event-AHU/OpenPAR
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。