arXiv:2601.14086cs.CVcs.AI2026-01

用双流Transformer融合图像与光流,提升动作识别准确率

Two-Stream temporal transformer for video action classification

论文配图:Two-Stream temporal transformer for video action classification
图 1 · 摘自论文原文
  • 双流架构分别处理视频帧与光流,捕捉时空特征
  • 在三个主流数据集上达到领先性能,验证模型有效性
  • 适合需要高精度动作识别的视觉系统开发

运动表征在视频理解中至关重要,广泛应用于动作识别、机器人导航等场景。近年来,基于自注意力机制的Transformer网络在诸多任务中表现出色。本文提出一种新型双流Transformer视频分类器,通过内容流(视频帧)和运动流(光流)提取时空信息。模型在联合光流与时间帧域中识别自注意力特征,并在Transformer编码器中建模其关系。实验结果表明,该方法在三个知名人体动作数据集上均取得优异分类效果。

原文摘要 · Abstract (English)

Motion representation plays an important role in video understanding and has many applications including action recognition, robot and autonomous guidance or others. Lately, transformer networks, through their self-attention mechanism capabilities, have proved their efficiency in many applications. In this study, we introduce a new two-stream transformer video classifier, which extracts spatio-temporal information from content and optical flow representing movement information. The proposed model identifies self-attention features across the joint optical flow and temporal frame domain and represents their relationships within the transformer encoder mechanism. The experimental results show that our proposed methodology provides excellent classification results on three well-known video datasets of human activities.

视频分类Transformer动作识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。