arXiv:2412.03093cs.CV2024-12被引 2

将CLIP模型适配到事件数据,实现跨模态零样本识别

Expanding Event Modality Applications through a Robust CLIP-Based Encoder

  • 用CLIP架构对齐事件与图像嵌入,支持零样本学习
  • 在物体识别任务中表现优异,无需额外训练即可泛化到视频事件
  • 集成于五模态框架,适用于图像、事件、文本、声音和深度数据

本文提出一种强大的编码器,将CLIP的能力迁移至事件数据,提升其在多领域的应用潜力。尽管大规模图像数据集推动了图像模型的发展,但事件数据集的匮乏限制了该模态的性能。为此,我们改造CLIP架构,使事件嵌入与图像嵌入对齐,支持零样本学习并保持文本对齐,同时缓解灾难性遗忘。该编码器在物体识别任务中表现强劲,在零样本与少样本学习中均达到竞争力水平。尤其值得注意的是,它能有效泛化至从视频中提取的事件数据,无需额外训练,展现出高度灵活性。此外,我们将该编码器整合进一个跨模态框架,支持图像、事件、文本、声音和深度五种模态的交互,拓展了跨模态应用的可能性。本工作凸显了稳健事件编码器的变革潜力,显著拓宽了事件数据在各领域中的应用范围。

原文摘要 · Abstract (English)

This paper introduces a powerful encoder that transfers CLIP`s capabilities to event-based data, enhancing its utility and expanding its applicability across diverse domains. While large-scale datasets have significantly advanced image-based models, the scarcity of comprehensive event datasets has limited performance potential in event modality. To address this challenge, we adapt CLIP`s architecture to align event embeddings with image embeddings, supporting zero-shot learning and preserving text alignment while mitigating catastrophic forgetting. Our encoder achieves strong performance in object recognition, with competitive results in zero-shot and few-shot learning tasks. Notably, it generalizes effectively to events extracted from video data without requiring additional training, highlighting its versatility. Additionally, we integrate this encoder within a cross-modality framework that facilitates interaction across five modalities-Image, Event, Text, Sound, and Depth-expanding the possibilities for cross-modal applications. Overall, this work underscores the transformative potential of a robust event encoder, broadening the scope and utility of event-based data across various fields.

事件数据CLIP跨模态零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。