用视频扩散模型生成真实运动数据,提升视频显著目标检测效果
TransFlow: Motion Knowledge Transfer from Video Diffusion Models to Video Salient Object Detection
- 从预训练视频扩散模型迁移语义化运动先验生成光流
- 在多个基准上实现性能提升,验证知识迁移有效性
- 适合需要高质量运动标注数据的研究者使用
视频显著目标检测(SOD)依赖运动线索区分显著物体与背景,但训练受限于稀少的视频数据集,远少于丰富的图像数据集。现有通过空间变换从静态图像生成视频序列的方法在运动引导任务中表现不佳,因其生成的光流不真实,缺乏对运动的语义理解。本文提出TransFlow,利用预训练视频扩散模型中的运动知识,为静态图像生成具有语义感知的光流。这些光流能体现物体在真实场景中的自然运动模式,同时保持空间边界和时间一致性。实验表明,该方法在多个基准上均取得性能提升,证明了运动知识的有效迁移。
原文摘要 · Abstract (English)
Video salient object detection (SOD) relies on motion cues to distinguish salient objects from backgrounds, but training such models is limited by scarce video datasets compared to abundant image datasets. Existing approaches that use spatial transformations to create video sequences from static images fail for motion-guided tasks, as these transformations produce unrealistic optical flows that lack semantic understanding of motion. We present TransFlow, which transfers motion knowledge from pre-trained video diffusion models to generate realistic training data for video SOD. Video diffusion models have learned rich semantic motion priors from large-scale video data, understanding how different objects naturally move in real scenes. TransFlow leverages this knowledge to generate semantically-aware optical flows from static images, where objects exhibit natural motion patterns while preserving spatial boundaries and temporal coherence. Our method achieves improved performance across multiple benchmarks, demonstrating effective motion knowledge transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。