用深度图生成伪光流,解决无监督视频目标分割数据不足问题
DepthFlow: Exploiting Depth-Flow Structural Correlations for Unsupervised Video Object Segmentation
- 从单张图像生成带结构信息的合成光流,保留关键视觉线索
- 在多个公开数据集上达到新最好性能,显著提升分割精度
- 适合研究视频分割、数据增强及无监督学习的学者参考
无监督视频对象分割(VOS)旨在识别视频中最具突出性的物体。近期,利用RGB图像和光流的双流方法受到广泛关注,但其性能受限于训练数据稀缺。为此,我们提出DepthFlow,一种新颖的数据生成方法,可从单张图像合成光流。该方法基于核心洞察:VOS模型更依赖于光流图中的结构信息而非几何准确性,且这种结构与深度高度相关。首先从源图像估计深度图,再将其转换为保留关键结构线索的合成光流场。该过程将大规模图像-掩码对转化为图像-光流-掩码训练对,极大扩展了网络训练数据。通过使用合成数据训练简单的编码器-解码器架构,我们在所有公开VOS基准测试中均取得新的最佳性能,验证了该方法在应对数据稀缺问题上的可扩展性与有效性。
原文摘要 · Abstract (English)
Unsupervised video object segmentation (VOS) aims to detect the most prominent object in a video. Recently, two-stream approaches that leverage both RGB images and optical flow have gained significant attention, but their performance is fundamentally constrained by the scarcity of training data. To address this, we propose DepthFlow, a novel data generation method that synthesizes optical flow from single images. Our approach is driven by the key insight that VOS models depend more on structural information embedded in flow maps than on their geometric accuracy, and that this structure is highly correlated with depth. We first estimate a depth map from a source image and then convert it into a synthetic flow field that preserves essential structural cues. This process enables the transformation of large-scale image-mask pairs into image-flow-mask training pairs, dramatically expanding the data available for network training. By training a simple encoder-decoder architecture with our synthesized data, we achieve new state-of-the-art performance on all public VOS benchmarks, demonstrating a scalable and effective solution to the data scarcity problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。