arXiv:2504.05830cs.CVcs.AI2025-04被引 2

融合RGB与事件相机,提升复杂场景下人体动作识别精度。

Human Activity Recognition using RGB-Event based Sensors: A Multi-modal Heat Conduction Model and A Benchmark Dataset

  • 基于热传导物理模型设计多模态特征融合框架
  • 构建包含107,646对视频的HARDVS 2.0大尺度数据集
  • 适用于低光、快速运动等挑战性场景的动作识别

人体动作识别(HAR)主要依赖传统RGB摄像头实现高性能识别,但在真实场景中,光照不足和快速运动等因素不可避免地降低其性能。为此,受生物启发的事件相机提供了一种克服传统RGB摄像头局限性的有前景方案。本文重新思考动作识别,结合RGB与事件相机数据。首个贡献是提出大规模多模态RGB-事件人体动作识别基准数据集HARDVS 2.0,填补了数据集空白。该数据集包含300类日常真实动作,共107,646对视频,覆盖多种挑战性场景。受物理启发的热传导模型启发,我们提出一种新颖的多模态热传导操作框架MMHCO-HAR。具体而言,先通过主干网络提取RGB帧与事件流特征嵌入,再设计多模态热传导模块进行双模态融合,核心为多模态热传导操作层。通过多模态DCT-IDCT层融合RGB与事件嵌入,并利用特征值嵌入(FVEs)自适应引入热导率系数。随后,提出基于策略路由的自适应融合模块以实现高性能分类。大量实验表明,本方法表现稳定,验证了其有效性和鲁棒性。源代码与数据集将发布于https://github.com/Event-AHU/HARDVS/tree/HARDVSv2。

原文摘要 · Abstract (English)

Human Activity Recognition (HAR) primarily relied on traditional RGB cameras to achieve high-performance activity recognition. However, the challenging factors in real-world scenarios, such as insufficient lighting and rapid movements, inevitably degrade the performance of RGB cameras. To address these challenges, biologically inspired event cameras offer a promising solution to overcome the limitations of traditional RGB cameras. In this work, we rethink human activity recognition by combining the RGB and event cameras. The first contribution is the proposed large-scale multi-modal RGB-Event human activity recognition benchmark dataset, termed HARDVS 2.0, which bridges the dataset gaps. It contains 300 categories of everyday real-world actions with a total of 107,646 paired videos covering various challenging scenarios. Inspired by the physics-informed heat conduction model, we propose a novel multi-modal heat conduction operation framework for effective activity recognition, termed MMHCO-HAR. More in detail, given the RGB frames and event streams, we first extract the feature embeddings using a stem network. Then, multi-modal Heat Conduction blocks are designed to fuse the dual features, the key module of which is the multi-modal Heat Conduction Operation layer. We integrate RGB and event embeddings through a multi-modal DCT-IDCT layer while adaptively incorporating the thermal conductivity coefficient via FVEs into this module. After that, we propose an adaptive fusion module based on a policy routing strategy for high-performance classification. Comprehensive experiments demonstrate that our method consistently performs well, validating its effectiveness and robustness. The source code and benchmark dataset will be released on https://github.com/Event-AHU/HARDVS/tree/HARDVSv2

动作识别多模态融合事件相机数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。