提出视频中人体与物体交互的时空检测新任务,精准识别交互行为和运动轨迹。
Spatial-Temporal Human-Object Interaction Detection
- 分两阶段:先追踪物体轨迹,再推理交互关系
- 在包含10,831个实例的数据集上超越现有方法
- 适合研究视频理解、人机交互与动作分析的学者
本文提出一个全新的视频级实例级人体-物体交互检测任务ST-HOID,旨在区分细粒度的人体-物体交互(HOIs)及主体与物体的时空轨迹。该任务源于人体-物体交互对以人为核心的视频内容理解至关重要。为解决此问题,我们设计了一种新方法,包含物体轨迹检测模块和交互推理模块。此外,我们构建了首个用于评估ST-HOID的基准数据集VidOR-HOID,包含10,831个时空交互实例。通过大量实验验证,所提方法在性能上优于当前最先进的图像级交互检测、视频视觉关系检测及视频人体-物体交互识别方法。
原文摘要 · Abstract (English)
In this paper, we propose a new instance-level human-object interaction detection task on videos called ST-HOID, which aims to distinguish fine-grained human-object interactions (HOIs) and the trajectories of subjects and objects. It is motivated by the fact that HOI is crucial for human-centric video content understanding. To solve ST-HOID, we propose a novel method consisting of an object trajectory detection module and an interaction reasoning module. Furthermore, we construct the first dataset named VidOR-HOID for ST-HOID evaluation, which contains 10,831 spatial-temporal HOI instances. We conduct extensive experiments to evaluate the effectiveness of our method. The experimental results demonstrate that our method outperforms the baselines generated by the state-of-the-art methods of image human-object interaction detection, video visual relation detection and video human-object interaction recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。