利用动作层次与上下文信息提升动作识别准确率。
Enhancing Action Recognition by Leveraging the Hierarchical Structure of Actions and Textual Context
- 设计专用Transformer融合视觉与文本上下文特征
- 在多个数据集上实现超过17%的准确率提升
- 适合老年人居家活动监测场景
我们提出一种新方法,通过利用动作的层次结构并引入包含位置和前序动作的上下文化文本信息,以反映动作的时间上下文。为此,我们设计了一种专用于动作识别的Transformer架构,同时使用来自RGB和光流数据的视觉特征以及代表上下文信息的文本嵌入。此外,我们定义了一个联合损失函数,以同时训练模型进行粗粒度和细粒度的动作识别,有效利用动作的层次特性。为验证方法有效性,我们扩展了Toyota Smarthome Untrimmed(TSU)数据集,加入动作层次结构,形成面向居家老人活动监测的层级化数据集Hierarchical TSU。消融实验评估了不同上下文与层次数据融合策略的影响。实验结果表明,该方法在Hierarchical TSU、Assembly101和IkeaASM数据集上持续优于现有最先进方法,顶1准确率提升超过17%。
原文摘要 · Abstract (English)
We propose a novel approach to improve action recognition by exploiting the hierarchical organization of actions and by incorporating contextualized textual information, including location and previous actions, to reflect the action's temporal context. To achieve this, we introduce a transformer architecture tailored for action recognition that employs both visual and textual features. Visual features are obtained from RGB and optical flow data, while text embeddings represent contextual information. Furthermore, we define a joint loss function to simultaneously train the model for both coarse- and fine-grained action recognition, effectively exploiting the hierarchical nature of actions. To demonstrate the effectiveness of our method, we extend the Toyota Smarthome Untrimmed (TSU) dataset by incorporating action hierarchies, resulting in the Hierarchical TSU dataset, a hierarchical dataset designed for monitoring activities of the elderly in home environments. An ablation study assesses the performance impact of different strategies for integrating contextual and hierarchical data. Experimental results demonstrate that the proposed method consistently outperforms SOTA methods on the Hierarchical TSU dataset, Assembly101 and IkeaASM, achieving over a 17% improvement in top-1 accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。