arXiv:2409.11441cs.CVcs.LG2024-09被引 1

让神经网络自主学习多层级运动流,持续生成一致的视觉特征。

Continual Learning of Conjugated Visual Representations through Higher-order Motion Flows

  • 通过自监督对比损失,让网络自主学习多层级运动流
  • 在真实与合成视频上显著优于现有无监督模型
  • 适合研究持续学习与运动感知表示的开发者

从连续视觉信息流中学习面临非独立同分布数据的挑战,但也提供了与信息流动保持一致的表征新机遇。本文研究了在多种运动诱导约束下进行无监督持续学习像素级特征的方法,称为运动共轭特征表示。不同于现有方法将运动视为给定信号(如真值或外部模块估计),本文让运动成为多层次特征层级中逐步自主学习的产物。通过神经网络估计多种运动流,涵盖从传统光流到由高层特征衍生的潜在信号,统称高阶运动。为防止持续学习导致平凡解,提出一种基于运动诱导相似性的空间感知自监督对比损失。在逼真合成数据流和真实视频上评估模型,相比预训练的先进特征提取器(基于Transformer)及近期无监督学习模型,性能显著更优。

原文摘要 · Abstract (English)

Learning with neural networks from a continuous stream of visual information presents several challenges due to the non-i.i.d. nature of the data. However, it also offers novel opportunities to develop representations that are consistent with the information flow. In this paper we investigate the case of unsupervised continual learning of pixel-wise features subject to multiple motion-induced constraints, therefore named motion-conjugated feature representations. Differently from existing approaches, motion is not a given signal (either ground-truth or estimated by external modules), but is the outcome of a progressive and autonomous learning process, occurring at various levels of the feature hierarchy. Multiple motion flows are estimated with neural networks and characterized by different levels of abstractions, spanning from traditional optical flow to other latent signals originating from higher-level features, hence called higher-order motions. Continuously learning to develop consistent multi-order flows and representations is prone to trivial solutions, which we counteract by introducing a self-supervised contrastive loss, spatially-aware and based on flow-induced similarity. We assess our model on photorealistic synthetic streams and real-world videos, comparing to pre-trained state-of-the art feature extractors (also based on Transformers) and to recent unsupervised learning models, significantly outperforming these alternatives.

持续学习运动感知自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。