arXiv:2504.08012cs.CV2025-04CVPR被引 2

用注意力融合时空相关性,提升视频预测的细节保真度。

SRVP: Strong Recollection Video Prediction Model Using Attention-Based Spatiotemporal Correlation Fusion

  • 引入标准与强化特征注意力模块,融合时空关联信息。
  • 在三个基准数据集上显著减少图像质量退化,性能媲美无RNN架构。
  • 适合关注视频生成细节还原的研究者与工程师。

视频预测(VP)通过利用过去帧的空间表征和时间上下文生成未来帧。传统基于循环神经网络(RNN)的模型通过增强记忆单元结构来捕捉长时间的时空状态,但存在物体外观细节逐渐丢失的问题。为解决该问题,本文提出强回忆视频预测(SRVP)模型,集成标准注意力(SA)与强化特征注意力(RFA)模块。两者均采用缩放点积注意力提取时间上下文与空间相关性,并进行融合以增强时空表征。在三个基准数据集上的实验表明,SRVP缓解了RNN模型的图像质量退化,同时达到与无RNN架构相当的预测性能。

原文摘要 · Abstract (English)

Video prediction (VP) generates future frames by leveraging spatial representations and temporal context from past frames. Traditional recurrent neural network (RNN)-based models enhance memory cell structures to capture spatiotemporal states over extended durations but suffer from gradual loss of object appearance details. To address this issue, we propose the strong recollection VP (SRVP) model, which integrates standard attention (SA) and reinforced feature attention (RFA) modules. Both modules employ scaled dot-product attention to extract temporal context and spatial correlations, which are then fused to enhance spatiotemporal representations. Experiments on three benchmark datasets demonstrate that SRVP mitigates image quality degradation in RNN-based models while achieving predictive performance comparable to RNN-free architectures.

视频预测注意力机制时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。