用强化学习动态选视频关键帧,提升预训练效率与动作识别性能
Reinforcement Learning meets Masked Video Modeling : Trajectory-Guided Adaptive Token Selection
- 基于运动轨迹自适应选择视频令牌,替代固定掩码策略
- 在4个数据集上实现高遮蔽率下仍保持优异动作识别精度
- 适合追求高效视频预训练的模型开发者与研究者
掩码视频建模(MVM)已成为视觉基础模型的有效预训练策略,通过利用可见时空令牌信息重建被遮掩的令牌。然而,其核心挑战在于如何选择合适的掩码策略。以往方法采用随机或管状掩码,或依赖外部预训练模型提供的运动先验、光流和语义线索。本文提出一种通用的轨迹感知自适应令牌采样器(TATS),建模令牌运动动态,可无缝集成至掩码自编码器(MAE)框架中,实现以运动为中心的令牌选择。同时,提出统一训练策略,使用近端策略优化(PPO)从零开始联合优化MAE与TATS。实验表明,该模型可在不降低动作识别下游任务性能的前提下实现高遮蔽率,且预训练过程内存效率高。在Something-Something v2、Kinetics-400、UCF101和HMDB51四个基准上的大量实验验证了方法的有效性、迁移性、泛化性和效率优于当前最优方法。
原文摘要 · Abstract (English)
Masked video modeling~(MVM) has emerged as a highly effective pre-training strategy for visual foundation models, whereby the model reconstructs masked spatiotemporal tokens using information from visible tokens. However, a key challenge in such approaches lies in selecting an appropriate masking strategy. Previous studies have explored predefined masking techniques, including random and tube-based masking, as well as approaches that leverage key motion priors, optical flow and semantic cues from externally pre-trained models. In this work, we introduce a novel and generalizable Trajectory-Aware Adaptive Token Sampler (TATS), which models the motion dynamics of tokens and can be seamlessly integrated into the masked autoencoder (MAE) framework to select motion-centric tokens in videos. Additionally, we propose a unified training strategy that enables joint optimization of both MAE and TATS from scratch using Proximal Policy Optimization (PPO). We show that our model allows for aggressive masking without compromising performance on the downstream task of action recognition while also ensuring that the pre-training remains memory efficient. Extensive experiments of the proposed approach across four benchmarks, including Something-Something v2, Kinetics-400, UCF101, and HMDB51, demonstrate the effectiveness, transferability, generalization, and efficiency of our work compared to other state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。