提出可被大模型理解的事件视觉表示方法,提升零样本识别性能。
LLM-EvRep: Learning an LLM-Compatible Event Representation Using a Self-Supervised Framework
- 自监督训练生成适配大模型的事件表示
- 在三个数据集上比E2VID提升最高50.21%
- 适合研究大模型与事件视觉融合的学者
事件驱动视觉识别近年进展显著,但多数方法依赖大量训练,难以高效处理事件数据。与此同时,大语言模型(LLMs)在多个领域展现出强大零样本能力,但在事件视觉识别中的应用仍较少。为弥合这一差距,我们提出 extbf{LLM-EvGen},一个事件表示生成器,用于生成适配大模型的事件表示 extbf{LLM-EvRep},从而提升大模型在事件识别任务中的表现。该生成器采用自监督框架训练,确保生成表示在语义一致性和结构保真度上均达标。在N-ImageNet、N-Caltech101和N-MNIST三个数据集上的实验表明,使用GPT-4o评估时,本方法相比事件转视频方法E2VID,在识别任务中分别提升15.93%、0.82%和50.21%。
原文摘要 · Abstract (English)
Recent advancements in event-based recognition have demonstrated significant promise, yet most existing approaches rely on extensive training, limiting their adaptability for efficient processing of event-driven visual content. Meanwhile, large language models (LLMs) have exhibited remarkable zero-shot capabilities across diverse domains, but their application to event-based visual recognition remains largely unexplored. To bridge this gap, we propose \textbf{LLM-EvGen}, an event representation generator that produces LLM-compatible event representations \textbf{LLM-EvRep}, thereby enhancing the performance of LLMs on event recognition tasks. The generator is trained using a self-supervised framework, aligning the generated representations with semantic consistency and structural fidelity. Comprehensive experiments were conducted on three datasets: N-ImageNet, N-Caltech101, and N-MNIST. The results demonstrate that our method, \textbf{LLM-EvRep}, outperforms the event-to-video method, E2VID, by 15.93\%, 0.82\%, and 50.21\%, respectively, in recognition tasks when evaluated using GPT-4o.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。