arXiv:2607.10470cs.CVeess.IV2026-07中稿 · ECCV

测试了光流模型在真实视频中的表现,发现合成数据上的成绩不靠谱。

On the Real-World Generalisability of Optical Flow Models

论文配图:On the Real-World Generalisability of Optical Flow Models
图 1 · 摘自论文原文
  • 构建真实世界评估基准FlowFactor,包含8204对真实视频帧
  • 光照变化和大位移场景表现与真实性能相关性最强,小运动反而易出错
  • 单纯增加训练数据无法解决泛化差距,需创新方法而非盲目扩增

将视觉模型投入现实应用是研究的核心目标。然而,由于真实光流难以获取,研究长期依赖合成数据和特定领域基准。本文探究这种差距的严重性,考察现代光流模型在真实视频中的泛化能力,并质疑合成基准上的精度是否能预测真实表现。为此,我们构建了一个包含8,204帧对的真实世界评估基准,涵盖TAP-Flow、Slow Flow及自建的FlowFactor数据集。FlowFactor为人工标注的真实世界数据,含1,000对高清帧,按四大混淆因素组织:大位移、重复纹理、遮挡和光照变化,每种设置仅改变一个变量,支持诊断分析。结果表明,光照变化与大位移的表现与真实性能相关性最强;大运动性能提升可能牺牲小运动、静止场景的鲁棒性。实验显示,Sintel、KITTI和Spring上的进展对真实数据预测能力弱,凸显建立广泛真实世界基准的必要性。有趣的是,增加训练数据量并不能缓解差距,亟需创新研究而非单纯扩大数据与算力。

原文摘要 · Abstract (English)

Real-world deployment of vision models to broadly benefit society is arguably a main research objective. In optical flow, however, the difficulty to obtain the ground truth has focused research mainly on synthetic data and domain-specific benchmarks. Here, we investigate the severity of this mismatch. We study how well modern optical flow estimation models generalise to real-world video and question if accuracy on synthetic benchmark proxies actually predicts accuracy on real-world optical flow. To address this, we build a real-world evaluation benchmark and evaluate the real-world generalisability of a broad set of recent optical flow models using standard checkpoints. Our benchmark contains 8,204 frame pairs across TAP-Flow, Slow Flow, and our own dataset FlowFactor. FlowFactor is a manually annotated real-world benchmark of 1,000 HD frame pairs organised into four confounding factors: large displacements, repetitive textures, occlusions, and lighting variation. Each setting mainly varies only one factor, enabling diagnostic, confounder-specific analysis. Using FlowFactor, we reveal that performance on varying lighting and large displacements correlates most strongly with real-world accuracy, and that improvements on large-motion regimes can trade off against robustness in small-motion, stationary scenes. Our experiments show that progress on Sintel, KITTI and Spring only weakly predicts accuracy on real-world data, highlighting the need for a broad real-world optical flow benchmark. Interestingly, scaling up the amount of training data does not necessarily resolve the gap, calling for new innovative research instead of simply scaling data and compute.

光流估计真实世界泛化基准测试模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。