用重建方法连接事件相机与多模态大模型,提升极端光照下视觉理解能力。
Reconstruction as a Bridge for Event-Based Visual Question Answering
- 通过帧重建与分词桥梁,融合事件数据与帧模型优势。
- 在1000对真实事件问答数据上达到当前最优性能。
- 适合关注事件视觉与大模型融合的研究者。
将事件相机与多模态大语言模型(MLLMs)结合,有望在极端视觉条件下实现通用场景理解,但需权衡事件数据的独特优势与与基于帧模型的兼容性。本文提出以重建为桥梁,设计了简单的帧重建与分词(FRT)方法,并开发了利用事件稀疏性的高效自适应重建与分词(ART)方法。为确保评估可靠性,我们构建了首个面向事件基MLLMs的客观真实世界基准EvQA,包含来自22个公开数据集的1000组事件-问答对。实验表明,所提方法在EvQA上达到当前最佳性能,凸显了MLLM在事件视觉中的巨大潜力。
原文摘要 · Abstract (English)
Integrating event cameras with Multimodal Large Language Models (MLLMs) promises general scene understanding in challenging visual conditions, yet requires navigating a trade-off between preserving the unique advantages of event data and ensuring compatibility with frame-based models. We address this challenge by using reconstruction as a bridge, proposing a straightforward Frame-based Reconstruction and Tokenization (FRT) method and designing an efficient Adaptive Reconstruction and Tokenization (ART) method that leverages event sparsity. For robust evaluation, we introduce EvQA, the first objective, real-world benchmark for event-based MLLMs, comprising 1,000 event-Q&A pairs from 22 public datasets. Our experiments demonstrate that our methods achieve state-of-the-art performance on EvQA, highlighting the significant potential of MLLMs in event-based vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。