让大模型读懂长时间事件相机数据,实现高精度跨模态理解。
LET-US: Long Event-Text Understanding of Scenes
- 用自适应压缩和文本引导查询减少事件流数据量,保留关键视觉信息。
- 在长序列事件数据上超越现有模型,在描述准确性和语义理解上领先。
- 适合研究事件相机、多模态大模型或长时序视频理解的学者使用。
事件相机以微秒级时间分辨率输出稀疏异步的事件流,支持低延迟与高动态范围的视觉感知。尽管现有多模态大模型(MLLMs)在理解RGB视频方面取得显著进展,但它们或难以有效解析事件流,或仅限于极短序列。本文提出LET-US框架,用于长事件流-文本理解,采用自适应压缩机制降低输入事件体积,同时保留关键视觉细节。为弥合事件流与文本表征间的巨大模态差距,我们采用两阶段优化策略,逐步提升模型对事件场景的解读能力。针对长事件流中的海量时间信息,利用文本引导的跨模态查询进行特征压缩,并结合分层聚类与相似性计算,提炼最具代表性的事件特征。此外,我们构建了大规模事件-文本对齐数据集以训练模型,实现了事件特征在大语言模型嵌入空间中的更紧密对齐。我们还设计了一个涵盖推理、描述生成、分类、时间定位与瞬间检索的综合性基准。实验表明,LET-US在长时事件流上的描述准确率与语义理解能力均优于现有最先进方法。所有数据集、代码与模型将公开共享。
原文摘要 · Abstract (English)
Event cameras output event streams as sparse, asynchronous data with microsecond-level temporal resolution, enabling visual perception with low latency and a high dynamic range. While existing Multimodal Large Language Models (MLLMs) have achieved significant success in understanding and analyzing RGB video content, they either fail to interpret event streams effectively or remain constrained to very short sequences. In this paper, we introduce LET-US, a framework for long event-stream--text comprehension that employs an adaptive compression mechanism to reduce the volume of input events while preserving critical visual details. LET-US thus establishes a new frontier in cross-modal inferential understanding over extended event sequences. To bridge the substantial modality gap between event streams and textual representations, we adopt a two-stage optimization paradigm that progressively equips our model with the capacity to interpret event-based scenes. To handle the voluminous temporal information inherent in long event streams, we leverage text-guided cross-modal queries for feature reduction, augmented by hierarchical clustering and similarity computation to distill the most representative event features. Moreover, we curate and construct a large-scale event-text aligned dataset to train our model, achieving tighter alignment of event features within the LLM embedding space. We also develop a comprehensive benchmark covering a diverse set of tasks -- reasoning, captioning, classification, temporal localization and moment retrieval. Experimental results demonstrate that LET-US outperforms prior state-of-the-art MLLMs in both descriptive accuracy and semantic comprehension on long-duration event streams. All datasets, codes, and models will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。