用混合光流补偿动态区域预测误差,提升视频生成真实感。
Lightweight Stochastic Video Prediction via Hybrid Warping
- 结合前向与后向光流融合,增强运动区域预测稳定性
- 在两个基准数据集上达到当前最优性能,有效应对运动不确定性
- 基于MobileNet的轻量化设计,适合实时应用
深度神经网络在动态区域的准确视频预测是计算机视觉中关键任务,尤其在自动驾驶、远程办公和远程医疗等场景中至关重要。由于固有的不确定性,现有预测模型常难以处理复杂的运动动态和遮挡问题。本文提出一种新型的随机长期视频预测模型,专注于动态区域,采用混合光流变形策略。通过整合前向和后向光流生成的帧,该方法有效弥补了单一技术的缺陷,提升了运动区域预测的准确性和真实性,并通过随机预测建模多种可能运动以应对不确定性。此外,为实现实时预测,引入基于MobileNet的轻量化架构。所提模型(SVPHW)在两个基准数据集上均取得当前最优表现。
原文摘要 · Abstract (English)
Accurate video prediction by deep neural networks, especially for dynamic regions, is a challenging task in computer vision for critical applications such as autonomous driving, remote working, and telemedicine. Due to inherent uncertainties, existing prediction models often struggle with the complexity of motion dynamics and occlusions. In this paper, we propose a novel stochastic long-term video prediction model that focuses on dynamic regions by employing a hybrid warping strategy. By integrating frames generated through forward and backward warpings, our approach effectively compensates for the weaknesses of each technique, improving the prediction accuracy and realism of moving regions in videos while also addressing uncertainty by making stochastic predictions that account for various motions. Furthermore, considering real-time predictions, we introduce a MobileNet-based lightweight architecture into our model. Our model, called SVPHW, achieves state-of-the-art performance on two benchmark datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。