arXiv:2511.18920cs.CV2025-11

用事件相机思路提升视频大模型效率,降耗提速还更准。

EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models

  • 通过事件触发机制,分层采样关键帧,剔除冗余视频内容。
  • 实现3.01倍计算量减少、3.10倍预填充加速,性能不降反升。
  • 适合追求高效视频理解的开发者,尤其关注推理成本优化者。

视频大语言模型虽具强大理解能力,但长视频带来的海量标记导致推理成本高昂。受事件视觉启发,我们提出无需训练的事件引导型高效时空理解框架EventSTU。时间维度上,设计粗到细的关键帧采样算法,利用事件相机的触发特性剔除冗余帧;空间维度上,基于事件视觉显著性设计自适应标记剪枝算法,实现零成本先验引导的空间压缩。从整体时空视角出发,进一步融合问题相关性动态分配标记剪枝预算。为评估构建了首个包含事件数据的人工标注多模态基准EventBench,覆盖多样化真实场景。除真实事件相机外,EventSTU也支持通过模拟事件进行通用视频理解。大量实验表明,相较于最强基线,EventSTU实现3.01倍FLOPs减少与3.10倍预填充速度提升,同时性能仍有所提高。

原文摘要 · Abstract (English)

Video large language models have demonstrated strong video understanding capabilities but suffer from high inference costs due to the massive number of tokens in long videos. Inspired by event-based vision, we propose an event-guided, training-free framework for efficient spatio-temporal understanding, named EventSTU. In the temporal domain, we design a coarse-to-fine keyframe sampling algorithm that exploits the change-triggered property of event cameras to eliminate redundant frames. In the spatial domain, we design an adaptive token pruning algorithm that leverages the visual saliency of events as a zero-cost prior to guide spatial reduction. From a holistic spatio-temporal perspective, we further integrate question relevance from keyframe sampling to adaptively allocate token pruning budgets. To facilitate evaluation, we construct EventBench, the first event-inclusive, human-annotated multimodal benchmark that covers diverse real-world scenarios. Beyond physical event cameras, EventSTU also supports general video understanding using simulated events. Comprehensive experiments show that EventSTU achieves 3.01x FLOPs reduction and 3.10x prefilling speedup over the strongest baseline while still improving performance.

视频理解事件视觉模型压缩高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。