arXiv:2409.12778cs.CV2024-09被引 2

用语言引导重构事件帧,实现无监督跨模态知识迁移

EventDance++: Language-guided Unsupervised Source-free Cross-modal Adaptation for Event-based Object Recognition

  • 通过语言引导重建事件到图像,构建代理图像
  • 在三个数据集上达到使用源数据方法的性能水平
  • 适合做事件视觉识别且无标签图像数据的研究者

本文解决事件视觉识别中无标签源图像数据下的跨模态(图像到事件)适应问题。由于图像与事件间存在巨大模态差异,仅依靠预训练源模型时,关键挑战在于从该模型中提取知识并有效迁移到事件域。受语言跨模态语义表达能力启发,我们提出 EventDance++,一种基于语言引导的无监督源自由跨模态适应框架。引入语言引导的重建式模态桥接(L-RMB)模块,以自监督方式从事件重建强度帧,并利用视觉-语言模型提供额外监督,丰富代理图像,增强模态桥接能力。由此生成的代理图像可用于从源模型中提取知识(即标签)。进一步提出多表示知识适配(MKA)模块,利用多种事件表示充分捕捉事件的时空特性,实现知识迁移。L-RMB 与 MKA 模块联合优化,有效缩小模态差距。在三个基准数据集上的实验表明,EventDance++ 性能媲美使用源数据的方法,验证了语言引导策略在事件视觉识别中的有效性。

原文摘要 · Abstract (English)

In this paper, we address the challenging problem of cross-modal (image-to-events) adaptation for event-based recognition without accessing any labeled source image data. This task is arduous due to the substantial modality gap between images and events. With only a pre-trained source model available, the key challenge lies in extracting knowledge from this model and effectively transferring knowledge to the event-based domain. Inspired by the natural ability of language to convey semantics across different modalities, we propose EventDance++, a novel framework that tackles this unsupervised source-free cross-modal adaptation problem from a language-guided perspective. We introduce a language-guided reconstruction-based modality bridging (L-RMB) module, which reconstructs intensity frames from events in a self-supervised manner. Importantly, it leverages a vision-language model to provide further supervision, enriching the surrogate images and enhancing modality bridging. This enables the creation of surrogate images to extract knowledge (i.e., labels) from the source model. On top, we propose a multi-representation knowledge adaptation (MKA) module to transfer knowledge to target models, utilizing multiple event representations to capture the spatiotemporal characteristics of events fully. The L-RMB and MKA modules are jointly optimized to achieve optimal performance in bridging the modality gap. Experiments on three benchmark datasets demonstrate that EventDance++ performs on par with methods that utilize source data, validating the effectiveness of our language-guided approach in event-based recognition.

事件视觉跨模态语言引导无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。