arXiv:2603.23032cs.CVcs.RO2026-03被引 2

用图像模型知识指导事件相机预训练,提升动态场景理解能力

Generative Event Pretraining with Foundation Model Alignment

  • 用图像语义对齐事件编码器,让事件数据获得语义基础
  • 通过混合事件-图像序列自回归预训练,捕捉事件独特时序特征
  • 在识别、分割、深度估计等任务上超越现有方法,适合高动态场景应用

事件相机凭借微秒级延迟和高动态范围,在快速运动和极端光照条件下提供鲁棒的视觉信号。然而,其独特的感知特性及标注数据有限,使得训练可迁移的事件视觉基础模型(VFMs)面临挑战。为此,我们提出GEP(生成式事件预训练),一种两阶段框架:首先,通过联合回归-对比目标将事件编码器与冻结的视觉基础模型对齐,使事件特征具备图像语义基础;其次,使用Transformer骨干网络在混合事件-图像序列上进行自回归预训练,以捕捉事件特有的时序结构。该方法在多种下游任务(包括物体识别、分割和深度估计)中优于现有最先进事件预训练方法。VFM引导的对齐与生成式序列建模相结合,构建出语义丰富、时序敏感的事件模型,可在不同领域间实现稳健泛化。

原文摘要 · Abstract (English)

Event cameras provide robust visual signals under fast motion and challenging illumination conditions thanks to their microsecond latency and high dynamic range. However, their unique sensing characteristics and limited labeled data make it challenging to train event-based visual foundation models (VFMs), which are crucial for learning visual features transferable across tasks. To tackle this problem, we propose GEP (Generative Event Pretraining), a two-stage framework that transfers semantic knowledge learned from internet-scale image datasets to event data while learning event-specific temporal dynamics. First, an event encoder is aligned to a frozen VFM through a joint regression-contrastive objective, grounding event features in image semantics. Second, a transformer backbone is autoregressively pretrained on mixed event-image sequences to capture the temporal structure unique to events. Our approach outperforms state-of-the-art event pretraining methods on a diverse range of downstream tasks, including object recognition, segmentation, and depth estimation. Together, VFM-guided alignment and generative sequence modeling yield a semantically rich, temporally aware event model that generalizes robustly across domains.

事件相机视觉基础模型自回归预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。