用多参考帧动态记忆提升视频目标追踪精度
Enhancing Self-Supervised Fine-Grained Video Object Tracking with Dynamic Memory Prediction
- 通过动态选择参考帧,直接利用多帧信息增强重建
- 在两个细粒度追踪任务上超越现有自监督方法
- 适合处理遮挡和快速运动等复杂场景
视频分析的成功依赖于跨帧像素的准确识别。基于视频对应关系学习的帧重建方法因其高效性而广受欢迎。然而,现有方法虽高效,却忽视了多参考帧在重建与决策中的直接作用,尤其在遮挡或快速运动等复杂情况下表现不足。本文提出动态记忆预测(DMP)框架,创新性地利用多参考帧直接增强帧重建。核心组件为参考帧记忆引擎,根据目标像素特征动态选择参考帧以提升追踪精度;同时构建双向目标预测网络,利用多参考帧提高模型鲁棒性。实验表明,该算法在两个细粒度视频目标追踪任务(物体分割与关键点追踪)上均优于当前最先进的自监督技术。
原文摘要 · Abstract (English)
Successful video analysis relies on accurate recognition of pixels across frames, and frame reconstruction methods based on video correspondence learning are popular due to their efficiency. Existing frame reconstruction methods, while efficient, neglect the value of direct involvement of multiple reference frames for reconstruction and decision-making aspects, especially in complex situations such as occlusion or fast movement. In this paper, we introduce a Dynamic Memory Prediction (DMP) framework that innovatively utilizes multiple reference frames to concisely and directly enhance frame reconstruction. Its core component is a Reference Frame Memory Engine that dynamically selects frames based on object pixel features to improve tracking accuracy. In addition, a Bidirectional Target Prediction Network is built to utilize multiple reference frames to improve the robustness of the model. Through experiments, our algorithm outperforms the state-of-the-art self-supervised techniques on two fine-grained video object tracking tasks: object segmentation and keypoint tracking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。