统一处理图像与事件数据,提升多模态感知能力
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework
- 融合RGB图像、事件图像和事件体素,实现多模态预训练
- 在5个下游任务中表现优异,显著提升多模态融合性能
- 适合视觉-事件融合场景,如自动驾驶与机器人感知
事件相机因其高动态范围、高时间分辨率、低功耗和低延迟等优势近年来受到广泛关注。一些研究开始探索直接在事件数据上进行预训练,但这些方法往往未能与RGB帧建立强关联,限制了其在多模态融合中的应用。为此,我们提出一种新的CM3AE多模态预训练框架,用于RGB-事件感知。该框架接受多种模态输入,包括RGB图像、事件图像和事件体素,为基于事件和RGB-事件融合的下游任务提供强大支持。具体而言,我们设计了一个多模态融合重建模块,从融合的多模态特征中重建原始图像,显式增强模型对跨模态互补信息的聚合能力。同时,采用多模态对比学习策略,将跨模态特征表示对齐至共享潜在空间,有效提升模型的多模态理解能力及全局依赖捕捉能力。我们构建了一个包含2,535,759对RGB-事件数据的大规模数据集用于预训练。在五个下游任务上的大量实验充分验证了CM3AE的有效性。源代码和预训练模型将发布于https://github.com/Event-AHU/CM3AE。
原文摘要 · Abstract (English)
Event cameras have attracted increasing attention in recent years due to their advantages in high dynamic range, high temporal resolution, low power consumption, and low latency. Some researchers have begun exploring pre-training directly on event data. Nevertheless, these efforts often fail to establish strong connections with RGB frames, limiting their applicability in multi-modal fusion scenarios. To address these issues, we propose a novel CM3AE pre-training framework for the RGB-Event perception. This framework accepts multi-modalities/views of data as input, including RGB images, event images, and event voxels, providing robust support for both event-based and RGB-event fusion based downstream tasks. Specifically, we design a multi-modal fusion reconstruction module that reconstructs the original image from fused multi-modal features, explicitly enhancing the model's ability to aggregate cross-modal complementary information. Additionally, we employ a multi-modal contrastive learning strategy to align cross-modal feature representations in a shared latent space, which effectively enhances the model's capability for multi-modal understanding and capturing global dependencies. We construct a large-scale dataset containing 2,535,759 RGB-Event data pairs for the pre-training. Extensive experiments on five downstream tasks fully demonstrated the effectiveness of CM3AE. Source code and pre-trained models will be released on https://github.com/Event-AHU/CM3AE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。