arXiv:2512.03837cs.CV2025-12TPAMI被引 3

用热图池化网络提升动作识别准确率

Heatmap Pooling Network for Action Recognition from RGB Videos

论文配图:Heatmap Pooling Network for Action Recognition from RGB Videos
图 1 · 摘自论文原文
  • 设计反馈式热图池化模块,提取紧凑鲁棒的体感特征
  • 在多个数据集上优于现有方法,最高提升1.8%准确率
  • 适合需要高效动作识别的智能监控与人机交互场景

RGB视频中的人类动作识别因信息丰富而受到广泛关注。然而,现有方法在提取深层特征时面临信息冗余、易受噪声干扰及存储成本高等问题。为此,我们提出一种新型热图池化网络(HP-Net),通过反馈池化模块提取视频中人体信息丰富、鲁棒且紧凑的聚合特征,显著优于先前的姿态数据与热图特征。此外,设计空间-运动协同学习模块与文本精炼调制模块,融合多模态数据以增强动作识别鲁棒性。在NTU RGB+D 60、NTU RGB+D 120、Toyota-Smarthome和UAV-Human等多个基准上进行大量实验,结果一致验证了HP-Net的有效性,其性能超越现有动作识别方法。代码已公开于:https://github.com/liujf69/HPNet-Action。

原文摘要 · Abstract (English)

Human action recognition (HAR) in videos has garnered widespread attention due to the rich information in RGB videos. Nevertheless, existing methods for extracting deep features from RGB videos face challenges such as information redundancy, susceptibility to noise and high storage costs. To address these issues and fully harness the useful information in videos, we propose a novel heatmap pooling network (HP-Net) for action recognition from videos, which extracts information-rich, robust and concise pooled features of the human body in videos through a feedback pooling module. The extracted pooled features demonstrate obvious performance advantages over the previously obtained pose data and heatmap features from videos. In addition, we design a spatial-motion co-learning module and a text refinement modulation module to integrate the extracted pooled features with other multimodal data, enabling more robust action recognition. Extensive experiments on several benchmarks namely NTU RGB+D 60, NTU RGB+D 120, Toyota-Smarthome and UAV-Human consistently verify the effectiveness of our HP-Net, which outperforms the existing human action recognition methods. Our code is publicly available at: https://github.com/liujf69/HPNet-Action.

动作识别热图池化多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。