arXiv:2409.09628cs.CVcs.AI2024-09被引 2

大模型无需训练即可识别事件相机图像,性能远超现有方法。

Can Large Language Models Grasp Event Signals? Exploring Pure Zero-Shot Event-based Recognition

  • 用提示工程激发大模型理解事件视觉内容,实现纯零样本识别。
  • GPT-4o在N-ImageNet上准确率比现有最优方法高五个数量级。
  • 首次验证大模型可直接处理事件数据,适合计算机视觉初学者参考。

近期事件相机的零样本物体识别取得显著进展,但这些方法依赖大量训练且受限于CLIP特性。本文首次探索大语言模型(LLM)对事件视觉内容的理解能力。实验表明,无需额外训练或微调,结合CLIP的LLM即可实现纯零样本事件识别。我们评估了GPT-4o、4turbo及两个开源模型对事件数据的直接识别能力,在三个基准数据集上系统测试其准确性。结果表明,经精心设计提示后,LLM显著提升事件零样本识别性能。尤其值得注意的是,GPT-4o在N-ImageNet上的识别准确率超越现有先进方法五个数量级。代码已开源。

原文摘要 · Abstract (English)

Recent advancements in event-based zero-shot object recognition have demonstrated promising results. However, these methods heavily depend on extensive training and are inherently constrained by the characteristics of CLIP. To the best of our knowledge, this research is the first study to explore the understanding capabilities of large language models (LLMs) for event-based visual content. We demonstrate that LLMs can achieve event-based object recognition without additional training or fine-tuning in conjunction with CLIP, effectively enabling pure zero-shot event-based recognition. Particularly, we evaluate the ability of GPT-4o / 4turbo and two other open-source LLMs to directly recognize event-based visual content. Extensive experiments are conducted across three benchmark datasets, systematically assessing the recognition accuracy of these models. The results show that LLMs, especially when enhanced with well-designed prompts, significantly improve event-based zero-shot recognition performance. Notably, GPT-4o outperforms the compared models and exceeds the recognition accuracy of state-of-the-art event-based zero-shot methods on N-ImageNet by five orders of magnitude. The implementation of this paper is available at \url{https://github.com/ChrisYu-Zz/Pure-event-based-recognition-based-LLM}.

事件相机大模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。