用超图补全稀疏事件流,提升视觉事件感知能力
EvRainDrop: HyperGraph-guided Completion for Effective Frame and Event Stream Aggregation
- 用超图关联时空事件,实现跨时空的事件补全
- 在多个任务上达到领先性能,显著缓解空间稀疏问题
- 支持多模态融合,适合事件相机与视觉融合场景
事件相机产生空间稀疏但时间密集的异步事件流。主流事件表示学习方法通常使用事件帧、体素或张量作为输入,尽管取得了显著进展,但仍难以解决由空间稀疏性引起的欠采样问题。本文提出一种新型超图引导的时空事件流补全机制,通过超图连接不同时刻和空间位置的事件标记,并利用上下文信息传递完成稀疏事件的补全。该方法可灵活将RGB标记作为超图节点,实现多模态超图驱动的信息补全。随后,通过自注意力机制聚合不同时步的超图节点信息,实现多模态特征的有效学习与融合。在单标签和多标签事件分类任务上的大量实验充分验证了所提框架的有效性。
原文摘要 · Abstract (English)
Event cameras produce asynchronous event streams that are spatially sparse yet temporally dense. Mainstream event representation learning algorithms typically use event frames, voxels, or tensors as input. Although these approaches have achieved notable progress, they struggle to address the undersampling problem caused by spatial sparsity. In this paper, we propose a novel hypergraph-guided spatio-temporal event stream completion mechanism, which connects event tokens across different times and spatial locations via hypergraphs and leverages contextual information message passing to complete these sparse events. The proposed method can flexibly incorporate RGB tokens as nodes in the hypergraph within this completion framework, enabling multi-modal hypergraph-based information completion. Subsequently, we aggregate hypergraph node information across different time steps through self-attention, enabling effective learning and fusion of multi-modal features. Extensive experiments on both single- and multi-label event classification tasks fully validated the effectiveness of our proposed framework. The source code of this paper will be released on https://github.com/Event-AHU/EvRainDrop.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。