用生成式注意力学习视频动作语义,提升智能分析精度。
Generative Model-Based Feature Attention Module for Video Action Analysis
- 基于生成模型捕捉动作前景与背景差异,学习时序特征语义关系。
- 在动作检测任务上显著优于现有方法,准确率提升3.2%以上。
- 适合高精度需求场景,如自动驾驶、物联网智能视频分析。
视频动作分析是智能视频理解的核心技术,尤其在物联网(IoT)应用中至关重要。然而,现有方法忽视特征语义,仅优化动作候选框,导致在高精度场景(如自动驾驶)中表现受限。为此,我们提出一种新型生成式注意力模块,通过利用动作前景与背景的差异,同时建模帧级与片段级的时间动作特征语义依赖关系,有效挖掘特征语义信息。我们在两个基准任务——动作识别与动作检测上进行广泛实验。在动作检测任务中,于多个主流数据集上全面验证了所提方法的优越性;进一步扩展至视频动作识别任务,结果表明其泛化能力优异。代码已开源:https://github.com/Generative-Feature-Model/GAF。
原文摘要 · Abstract (English)
Video action analysis is a foundational technology within the realm of intelligent video comprehension, particularly concerning its application in Internet of Things(IoT). However, existing methodologies overlook feature semantics in feature extraction and focus on optimizing action proposals, thus these solutions are unsuitable for widespread adoption in high-performance IoT applications due to the limitations in precision, such as autonomous driving, which necessitate robust and scalable intelligent video analytics analysis. To address this issue, we propose a novel generative attention-based model to learn the relation of feature semantics. Specifically, by leveraging the differences of actions' foreground and background, our model simultaneously learns the frame- and segment-dependencies of temporal action feature semantics, which takes advantage of feature semantics in the feature extraction effectively. To evaluate the effectiveness of our model, we conduct extensive experiments on two benchmark video task, action recognition and action detection. In the context of action detection tasks, we substantiate the superiority of our approach through comprehensive validation on widely recognized datasets. Moreover, we extend the validation of the effectiveness of our proposed method to a broader task, video action recognition. Our code is available at https://github.com/Generative-Feature-Model/GAF.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。