arXiv:2604.03590cs.CV2026-04中稿 · ABAW2026

用三维度信息增强骨骼数据,提升视频动作识别准确率

SBF: An Effective Representation to Augment Skeleton for Video-based Human Action Recognition

论文配图:SBF: An Effective Representation to Augment Skeleton for Video-based Human Action Recognition
图 1 · 摘自论文原文
  • 提出SBF表示法,融合关节尺度、人体轮廓和人物交互流信息
  • 在多个数据集上显著超越纯骨骼方法,精度提升明显
  • 无需额外标注,可直接集成到现有骨架识别流程中

当前基于视频的人体动作识别(HAR)多采用2D骨骼作为中间表示,但其在常见场景中仍表现不佳,主要因无法捕捉关节深度、人体轮廓及人与物体交互等关键信息。为此,本文提出一种新方法,在HAR流程中通过引入新的表示形式Scale-Body-Flow(SBF)来增强骨骼数据。SBF包含三个分量:由关节尺度提供的尺度图体积(含深度信息)、描绘人体轮廓的体图,以及由像素级光流计算得出的人物-物体交互流图。为预测SBF,本文进一步设计SFSNet——一种仅需骨骼和光流监督的新型分割网络,无需额外标注成本。在多个数据集上的大量实验表明,基于SBF与SFSNet的管道在保持与顶尖纯骨骼方法相当的紧凑性与效率的同时,实现了显著更高的动作识别准确率。

原文摘要 · Abstract (English)

Many modern video-based human action recognition (HAR) approaches use 2D skeleton as the intermediate representation in their prediction pipelines. Despite overall encouraging results, these approaches still struggle in many common scenes, mainly because the skeleton does not capture critical action-related information pertaining to the depth of the joints, contour of the human body, and interaction between the human and objects. To address this, we propose an effective approach to augment skeleton with a representation capturing action-related information in the pipeline of HAR. The representation, termed Scale-Body-Flow (SBF), consists of three distinct components, namely a scale map volume given by the scale (and hence depth information) of each joint, a body map outlining the human subject, and a flow map indicating human-object interaction given by pixel-wise optical flow values. To predict SBF, we further present SFSNet, a novel segmentation network supervised by the skeleton and optical flow without extra annotation overhead beyond the existing skeleton extraction. Extensive experiments across different datasets demonstrate that our pipeline based on SBF and SFSNet achieves significantly higher HAR accuracy with similar compactness and efficiency as compared with the state-of-the-art skeleton-only approaches.

动作识别骨骼增强视觉表征光流

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。