通过动态拆分事件帧捕捉运动线索,提升光照变化下的视频追踪性能。
Dynamic Subframe Splitting and Spatio-Temporal Motion Entangled Sparse Attention for RGB-E Tracking
- 动态拆分事件流为细粒度片段,保留时间信息
- 设计稀疏注意力机制增强时空特征交互,精度优于现有方法
- 适合快速运动或低光照场景的多模态追踪任务
基于事件的仿生相机以高时间分辨率和高动态范围异步捕捉动态场景,在光照劣化与快速运动条件下,具备融合事件与RGB数据的潜力。现有RGB-E追踪方法在融合前使用Transformer注意力机制建模事件特征,但通常将事件流聚合为单个事件帧,忽略了事件流中固有的时间信息。此外,传统注意力机制适用于密集语义特征,而对稀疏事件特征的处理仍需革新。本文提出一种动态事件子帧拆分策略,将事件流分解为更细粒度的事件簇,以捕获包含运动线索的时空特征。基于此,设计了一种基于事件的稀疏注意力机制,增强事件特征在时空维度上的交互能力。实验结果表明,该方法在FE240与COESOT数据集上均优于现有最先进方法,为事件数据处理提供了有效方案。
原文摘要 · Abstract (English)
Event-based bionic camera asynchronously captures dynamic scenes with high temporal resolution and high dynamic range, offering potential for the integration of events and RGB under conditions of illumination degradation and fast motion. Existing RGB-E tracking methods model event characteristics utilising attention mechanism of Transformer before integrating both modalities. Nevertheless, these methods involve aggregating the event stream into a single event frame, lacking the utilisation of the temporal information inherent in the event stream.Moreover, the traditional attention mechanism is well-suited for dense semantic features, while the attention mechanism for sparse event features require revolution. In this paper, we propose a dynamic event subframe splitting strategy to split the event stream into more fine-grained event clusters, aiming to capture spatio-temporal features that contain motion cues. Based on this, we design an event-based sparse attention mechanism to enhance the interaction of event features in temporal and spatial dimensions. The experimental results indicate that our method outperforms existing state-of-the-art methods on the FE240 and COESOT datasets, providing an effective processing manner for the event data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。