arXiv:2411.13975cs.CV2024-11

用生成模型将静态图转为带真实运动的视频,提升显著目标检测效果。

Transforming Static Images Using Generative Models for Video Salient Object Detection

论文配图:Transforming Static Images Using Generative Models for Video Salient Object Detection
图 1 · 摘自论文原文
  • 用扩散模型生成图像到视频的动态变换,保留语义一致性。
  • 生成的光流更真实,能捕捉物体独立运动特性。
  • 大幅提升训练数据质量,实现多个数据集上最佳性能。

在许多视频处理任务中,利用大规模图像数据集是常见策略,因为图像数据更丰富且有助于知识迁移。传统模拟视频的方法通常对静态图像施加空间变换(如仿射变换、样条扭曲)以生成序列,模拟时间演变。然而,在视频显著目标检测等依赖外观与运动线索的任务中,这些基础的图像转视频方法无法生成真实的光流,难以体现每个物体的独立运动属性。本研究证明,图像到视频的扩散模型能够生成真实动态变换,同时理解图像组件间的上下文关系。该能力使模型可生成合理光流,在保持语义完整性的同时反映场景元素的独立运动。通过这种方式增强单张图像,我们构建了大规模图像-光流配对数据,显著提升模型训练效果。该方法在所有公开基准数据集上均达到领先性能,超越现有方法。

原文摘要 · Abstract (English)

In many video processing tasks, leveraging large-scale image datasets is a common strategy, as image data is more abundant and facilitates comprehensive knowledge transfer. A typical approach for simulating video from static images involves applying spatial transformations, such as affine transformations and spline warping, to create sequences that mimic temporal progression. However, in tasks like video salient object detection, where both appearance and motion cues are critical, these basic image-to-video techniques fail to produce realistic optical flows that capture the independent motion properties of each object. In this study, we show that image-to-video diffusion models can generate realistic transformations of static images while understanding the contextual relationships between image components. This ability allows the model to generate plausible optical flows, preserving semantic integrity while reflecting the independent motion of scene elements. By augmenting individual images in this way, we create large-scale image-flow pairs that significantly enhance model training. Our approach achieves state-of-the-art performance across all public benchmark datasets, outperforming existing approaches.

视频生成扩散模型显著目标检测光流生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。