arXiv:2508.05507cs.CV2025-08被引 1

通过物理启发的自监督预训练,从稀疏噪声事件数据中挖掘边缘与纹理信息。

Revealing Latent Information: A Physics-inspired Self-supervised Pre-training Framework for Noisy and Sparse Events

  • 基于事件物理采样机制设计差分掩码建模,重建时序强度差图增强原始事件特征。
  • 在下游任务中,对象识别、语义分割和光流估计均超越现有最优方法。
  • 适合处理高动态范围、低光照等挑战性场景下的事件视觉任务。

事件相机是一种新型类脑视觉传感器,以高时间分辨率和宽动态范围记录数据,为复杂场景下的精确视觉表征带来新可能。然而,事件数据本身稀疏且噪声大,主要反映亮度变化,导致有效特征提取困难。为此,我们提出一种自监督预训练框架,充分揭示事件数据中的潜在信息,包括边缘信息与纹理线索。该框架包含三个阶段:差异引导的掩码建模,受事件物理采样过程启发,重建时序强度差图,从原始事件数据中提取增强信息;主干固定特征迁移,不更新主干网络,以保留掩码建模所学表示并稳定对比学习效果;聚焦式对比学习,更新整个模型,通过关注高价值区域提升语义区分能力。大量实验表明,该框架在多个下游任务中表现稳健,持续优于当前最优方法,涵盖对象识别、语义分割和光流估计。代码与数据集已公开于 https://github.com/BIT-Vision/EventPretrain。

原文摘要 · Abstract (English)

Event camera, a novel neuromorphic vision sensor, records data with high temporal resolution and wide dynamic range, offering new possibilities for accurate visual representation in challenging scenarios. However, event data is inherently sparse and noisy, mainly reflecting brightness changes, which complicates effective feature extraction. To address this, we propose a self-supervised pre-training framework to fully reveal latent information in event data, including edge information and texture cues. Our framework consists of three stages: Difference-guided Masked Modeling, inspired by the event physical sampling process, reconstructs temporal intensity difference maps to extract enhanced information from raw event data. Backbone-fixed Feature Transition contrasts event and image features without updating the backbone to preserve representations learned from masked modeling and stabilizing their effect on contrastive learning. Focus-aimed Contrastive Learning updates the entire model to improve semantic discrimination by focusing on high-value regions. Extensive experiments show our framework is robust and consistently outperforms state-of-the-art methods on various downstream tasks, including object recognition, semantic segmentation, and optical flow estimation. The code and dataset are available at https://github.com/BIT-Vision/EventPretrain.

事件相机自监督学习神经形态视觉预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。