arXiv:2506.07740cs.CV2025-06TPAMI被引 13

从单张真实图像生成大规规模拟光流数据,提升实际应用性能

Flow-Anything: Learning Real-World Optical Flow Estimation from Large-Scale Single-view Images

  • 用单目深度网络将单图转为3D,虚拟摄像机渲染光流与新视角
  • 自研模块实现动态物体建模,生成真实感光流数据集(FA-Flow)
  • 模型可迁移至多视频任务,优于现有无监督与合成数据训练方法

光流估计是计算机视觉关键领域,支撑各类视频任务。但现有训练依赖动画合成数据,导致真实场景泛化能力受限,且难以有效扩展数据规模。为此,本文提出Flow-Anything,一种从任意单视图真实图像中学习光流估计的大规模数据生成框架。首先,利用先进单目深度估计网络将单图转为3D表示,进而通过虚拟摄像机渲染光流和新视角图像;其次,设计独立于物体的体渲染模块与基于深度的修复模块,对3D表示中的动态物体进行建模。该流程可从大规模真实图像生成逼真训练数据,构建出首个基于真实图像的光流数据集:FA-Flow。首次证明从大规模真实图像生成训练数据的有效性,在多个基准上超越最先进的无监督及合成数据训练方法。所提模型具备基础模型潜力,显著提升多种下游视频任务性能。

原文摘要 · Abstract (English)

Optical flow estimation is a crucial subfield of computer vision, serving as a foundation for video tasks. However, the real-world robustness is limited by animated synthetic datasets for training. This introduces domain gaps when applied to real-world applications and limits the benefits of scaling up datasets. To address these challenges, we propose \textbf{Flow-Anything}, a large-scale data generation framework designed to learn optical flow estimation from any single-view images in the real world. We employ two effective steps to make data scaling-up promising. First, we convert a single-view image into a 3D representation using advanced monocular depth estimation networks. This allows us to render optical flow and novel view images under a virtual camera. Second, we develop an Object-Independent Volume Rendering module and a Depth-Aware Inpainting module to model the dynamic objects in the 3D representation. These two steps allow us to generate realistic datasets for training from large-scale single-view images, namely \textbf{FA-Flow Dataset}. For the first time, we demonstrate the benefits of generating optical flow training data from large-scale real-world images, outperforming the most advanced unsupervised methods and supervised methods on synthetic datasets. Moreover, our models serve as a foundation model and enhance the performance of various downstream video tasks.

光流估计3D重建数据生成真实世界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。