arXiv:2409.13091cs.CVcs.AI2024-09中稿 · the Human-inspired…

通过引入3D信息提升视频动作识别的可解释性与准确性

Interpretable Action Recognition on Hard to Classify Actions

  • 基于物体位置、手部动作和时空关系建模动作理解
  • 加入深度关系后,对相似动作的识别准确率显著提升
  • 适合关注可解释性视频分析的研究者和开发者

我们研究了一种类人可解释的视频理解模型。人类通过识别物体与部件之间的关键时空关系来理解复杂动作,例如物体进入容器开口。为此,我们构建了一个利用物体与手部位置及其运动来识别动作的模型。针对三个最易混淆的动作类别,发现缺乏3D信息是主要问题。为此,我们通过两种方式增强模型的3D感知:(1)微调先进目标检测模型以区分“Container”与“NotContainer”,将物体形状信息融入特征;(2)使用先进的深度估计模型提取各物体的深度值,并计算深度关系,扩展原有关系表示。这些3D改进在Something-Something-v2数据集中三个表面相似的“Putting”动作子集上进行评估。结果显示,容器检测器未提升性能,但引入深度关系显著改善了识别效果。

原文摘要 · Abstract (English)

We investigate a human-like interpretable model of video understanding. Humans recognise complex activities in video by recognising critical spatio-temporal relations among explicitly recognised objects and parts, for example, an object entering the aperture of a container. To mimic this we build on a model which uses positions of objects and hands, and their motions, to recognise the activity taking place. To improve this model we focussed on three of the most confused classes (for this model) and identified that the lack of 3D information was the major problem. To address this we extended our basic model by adding 3D awareness in two ways: (1) A state-of-the-art object detection model was fine-tuned to determine the difference between "Container" and "NotContainer" in order to integrate object shape information into the existing object features. (2) A state-of-the-art depth estimation model was used to extract depth values for individual objects and calculate depth relations to expand the existing relations used our interpretable model. These 3D extensions to our basic model were evaluated on a subset of three superficially similar "Putting" actions from the Something-Something-v2 dataset. The results showed that the container detector did not improve performance, but the addition of depth relations made a significant improvement to performance.

视频理解可解释性3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。