动态建模视频上下文,让每帧都参与关键识别
Making Every Frame Matter: Continuous Activity Recognition in Streaming Video via Adaptive Video Context Modeling
- 按需保留时空相关特征,剔除无关视觉信息
- 在边缘设备上实现30+帧/秒,准确率提升1.2%至79.7%
- 适合实时视频分析与大模型视频编码任务
视频活动识别在机器人和具身AI中日益重要。连续视频活动识别面临流式视频快速膨胀带来的挑战,其包含多尺度、未剪裁的活动。本文提出CARS系统,通过自适应视频上下文建模克服上述问题。该方法在时空维度上选择性保留与活动相关的特征。CARS包含两项核心设计:一是通过消除无关视觉特征实现活动空间特征提取,保持识别精度;二是引入活动感知状态更新,增强动态适应性,更好保留视频上下文以支持多尺度活动识别。CARS在典型边缘设备上运行速度超过30 FPS,准确率相比所有基线提升1.2%至79.7%。此外,我们探索将CARS作为大视频模型的视频编码器,实验表明其在分布内视频活动数据集上带来0.46分提升(5分制),在零样本视频活动数据集上提升1.19%至4%。
原文摘要 · Abstract (English)
Video activity recognition has become increasingly important in robots and embodied AI. Recognizing continuous video activities poses considerable challenges due to the fast expansion of streaming video, which contains multi-scale and untrimmed activities. We introduce a novel system, CARS, to overcome these issues through adaptive video context modeling. Adaptive video context modeling refers to selectively maintaining activity-related features in temporal and spatial dimensions. CARS has two key designs. The first is an activity spatial feature extraction by eliminating irrelevant visual features while maintaining recognition accuracy. The second is an activity-aware state update introducing dynamic adaptability to better preserve the video context for multi-scale activity recognition. Our CARS runs at speeds $>$30 FPS on typical edge devices and outperforms all baselines by 1.2\% to 79.7\% in accuracy. Moreover, we explore applying CARS to a large video model as a video encoder. Experimental results show that our CARS can result in a 0.46-point enhancement (on a 5-point scale) on the in-distribution video activity dataset, and an improvement ranging from 1.19\% to 4\% on zero-shot video activity datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。