提出高分辨率运动特征表示,提升密集预测任务性能。
FlowFeat: Pixel-Dense Embedding of Motion Profiles

- 通过自监督学习嵌入多种可能的运动轨迹,生成高分辨率特征图。
- 在视频分割、单目深度估计等任务中显著提升5个主流模型表现。
- 对光流误差不敏感,适合缺乏标注数据的密集视觉任务场景。
密集且通用的图像表征是几乎所有计算机视觉应用成功的关键。然而,当前最先进的网络(如Transformer)生成的特征图分辨率较低,不利于密集预测任务。为此,我们提出FlowFeat,一种高分辨率、多任务的特征表示方法。其核心是新颖的蒸馏技术,用于嵌入一系列合理的表观运动分布(即运动轮廓)。通过利用光学流网络和多样化的视频数据,我们构建了一个有效的自监督训练框架,统计近似表观运动。凭借出色的局部空间细节,FlowFeat不仅编码了丰富的几何与语义线索,还展现出良好的时间一致性。实验表明,FlowFeat显著增强了五个前沿编码器及多种上采样策略在三项密集任务(视频目标分割、单目深度估计、语义分割)上的表征能力。训练流程计算开销低,对光流估计不准确具有鲁棒性,即使使用无监督光流网络仍保持高效。本工作推动了可靠且通用的密集图像表征的发展。
原文摘要 · Abstract (English)
Dense and versatile image representations underpin the success of virtually all computer vision applications. However, state-of-the-art networks, such as transformers, produce low-resolution feature grids, which are suboptimal for dense prediction tasks. To address this limitation, we present FlowFeat, a high-resolution and multi-task feature representation. The key ingredient behind FlowFeat is a novel distillation technique that embeds a distribution of plausible apparent motions, or motion profiles. By leveraging optical flow networks and diverse video data, we develop an effective self-supervised training framework that statistically approximates the apparent motion. With its remarkable level of spatial detail, FlowFeat encodes a compelling degree of geometric and semantic cues while exhibiting high temporal consistency. Empirically, FlowFeat significantly enhances the representational power of five state-of-the-art encoders and alternative upsampling strategies across three dense tasks: video object segmentation, monocular depth estimation and semantic segmentation. Training FlowFeat is computationally inexpensive and robust to inaccurate flow estimation, remaining highly effective even when using unsupervised flow networks. Our work takes a step forward towards reliable and versatile dense image representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。