用扰动机制提升事件相机光流估计的时空建模能力
Perturbed State Space Feature Encoders for Optical Flow with Event Cameras
- 通过扰动状态矩阵增强状态空间模型的时序建模能力
- 在DSEC-Flow和MVSEC上分别提升8.48%和11.86%的EPE性能
- 适合做事件相机光流任务的研究者与工业应用开发者
基于事件的相机因其对运动的响应特性,在光流估计方面相较于传统相机具有显著优势。尽管深度学习已超越传统方法,但现有用于事件相机的神经网络仍存在时空推理能力不足的问题。本文提出扰动状态空间特征编码器(P-SSE),用于多帧光流估计,以解决上述挑战。P-SSE通过大型感受野自适应处理时空特征,其计算复杂度保持线性,类似状态空间模型(SSM);而关键创新在于对控制状态动态的矩阵施加扰动,显著提升了模型稳定性和性能。我们将P-SSE集成至利用双向光流与递归连接的框架中,扩大了光流预测的时序上下文。在DSEC-Flow和MVSEC数据集上的评估表明,本模型分别实现8.48%和11.86%的EPE性能提升。
原文摘要 · Abstract (English)
With their motion-responsive nature, event-based cameras offer significant advantages over traditional cameras for optical flow estimation. While deep learning has improved upon traditional methods, current neural networks adopted for event-based optical flow still face temporal and spatial reasoning limitations. We propose Perturbed State Space Feature Encoders (P-SSE) for multi-frame optical flow with event cameras to address these challenges. P-SSE adaptively processes spatiotemporal features with a large receptive field akin to Transformer-based methods, while maintaining the linear computational complexity characteristic of SSMs. However, the key innovation that enables the state-of-the-art performance of our model lies in our perturbation technique applied to the state dynamics matrix governing the SSM system. This approach significantly improves the stability and performance of our model. We integrate P-SSE into a framework that leverages bi-directional flows and recurrent connections, expanding the temporal context of flow prediction. Evaluations on DSEC-Flow and MVSEC datasets showcase P-SSE's superiority, with 8.48% and 11.86% improvements in EPE performance, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。