arXiv:2412.07080cs.CVcs.AI2024-12被引 21

用自监督学习提升事件相机数据的表示质量,效果更好且通用性强。

EvRepSL: Event-Stream Representation via Self-Supervised Learning for Event-Based Vision

  • 基于时空统计构建事件流表示EvRep,理论推导其与视频帧关系
  • 通过自监督训练生成器RepGen,输出高质量事件表示EvRepSL
  • 无需微调,适配多种相机和任务,性能超越现有方法

事件流表示是使用事件相机进行计算机视觉任务的第一步,将异步事件流转化为可被传统机器学习模型处理的结构。然而,当前主流事件流表示多为人工设计,其质量难以保证,因事件流本身具有噪声特性。本文提出一种数据驱动方法,旨在提升事件流表示的质量。首先引入一种基于时空统计的新表示EvRep;随后理论上推导了异步事件流与同步视频帧之间的内在关联;基于此,以EvRep为输入,采用自监督学习方式训练表示生成器RepGen;最终通过训练好的RepGen将事件流转换为高质量表示EvRepSL(无需微调或重训练)。该方法在多种主流事件相机采集的分类与光流数据集上进行了充分验证,实验结果表明,本方法不仅显著优于现有事件流表示,且对不同事件相机和任务具有高度适应性。

原文摘要 · Abstract (English)

Event-stream representation is the first step for many computer vision tasks using event cameras. It converts the asynchronous event-streams into a formatted structure so that conventional machine learning models can be applied easily. However, most of the state-of-the-art event-stream representations are manually designed and the quality of these representations cannot be guaranteed due to the noisy nature of event-streams. In this paper, we introduce a data-driven approach aiming at enhancing the quality of event-stream representations. Our approach commences with the introduction of a new event-stream representation based on spatial-temporal statistics, denoted as EvRep. Subsequently, we theoretically derive the intrinsic relationship between asynchronous event-streams and synchronous video frames. Building upon this theoretical relationship, we train a representation generator, RepGen, in a self-supervised learning manner accepting EvRep as input. Finally, the event-streams are converted to high-quality representations, termed as EvRepSL, by going through the learned RepGen (without the need of fine-tuning or retraining). Our methodology is rigorously validated through extensive evaluations on a variety of mainstream event-based classification and optical flow datasets (captured with various types of event cameras). The experimental results highlight not only our approach's superior performance over existing event-stream representations but also its versatility, being agnostic to different event cameras and tasks.

事件相机自监督学习表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。