arXiv:2503.18933cs.CV2025-03

用多模态数据提升视频预测精度,支持单模态下仍保持高性能。

SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction

  • 基于预训练扩散模型,通过时空交叉注意力融合多模态信息。
  • 在Cityscapes、BAIR等数据集上达到当前最优效果,单模态也有效。
  • 适用于深度、语义、气候等多种模态,通用性强。

视频未来帧预测对决策系统至关重要,但仅依赖RGB图像常无法充分捕捉真实世界复杂性。为此,我们提出同步多模态视频预测框架SyncVP,融合互补模态数据以增强预测的丰富性与准确性。SyncVP基于预训练的模态专用扩散模型,并引入高效的时空交叉注意力模块,实现跨模态有效信息共享。我们在标准基准数据集(如Cityscapes和BAIR)上评估,以深度图为附加模态;同时在SYNTHIA(语义信息)和ERA5-Land(气候数据)上验证其泛化能力。结果表明,SyncVP在多种场景下均达当前最优性能,即使仅使用单一模态也表现稳健,展现出广泛的应用潜力。

原文摘要 · Abstract (English)

Predicting future video frames is essential for decision-making systems, yet RGB frames alone often lack the information needed to fully capture the underlying complexities of the real world. To address this limitation, we propose a multi-modal framework for Synchronous Video Prediction (SyncVP) that incorporates complementary data modalities, enhancing the richness and accuracy of future predictions. SyncVP builds on pre-trained modality-specific diffusion models and introduces an efficient spatio-temporal cross-attention module to enable effective information sharing across modalities. We evaluate SyncVP on standard benchmark datasets, such as Cityscapes and BAIR, using depth as an additional modality. We furthermore demonstrate its generalization to other modalities on SYNTHIA with semantic information and ERA5-Land with climate data. Notably, SyncVP achieves state-of-the-art performance, even in scenarios where only one modality is present, demonstrating its robustness and potential for a wide range of applications.

视频预测多模态扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。