用脉冲神经网络提升事件相机的特征检测精度与能效。
E-S2Feat:Semantic-Guided Spiking Local Feature Detection and Description for Event Cameras

- 设计模块化脉冲激活机制,低比特下保留细节特征。
- 引入语义引导调制,提升关键点分布稳定性和描述子区分度。
- 在资源受限设备上实现高精度与4.8倍能效提升,适合无人机等场景。
得益于高时间分辨率和动态范围,基于事件的局部特征方法受到越来越多关注。然而,事件稀疏性、噪声以及纹理有限仍阻碍鲁棒的局部特征学习。在无人机等资源受限平台部署时,还需兼顾精度与能效。为此,本文提出E-S2Feat,一种用于事件相机的脉冲神经网络框架,实现局部特征检测与描述的联合优化。首先,模块化脉冲激活机制在低比特、低功耗推理下保留细粒度结构线索与判别信息,提升整体表征保真度。其次,语义引导特征调制机制利用语义先验优化关键点响应分布,增强局部描述子的判别能力,引导模型提取几何更稳定、判别力更强的局部特征。在ECD与EDS数据集上的实验表明,该方法显著优于SuperEvent等基线方法,在姿态估计精度上表现优异。其精度接近人工神经网络对手,同时理论计算能耗降低约4.8倍。TUM-VIE数据集上的视觉惯性里程计实验进一步验证了该方法在完整SLAM系统中的有效性与实际应用潜力。
原文摘要 · Abstract (English)
Benefiting from high temporal resolution and dynamic range, event-based local feature methods have attracted increasing attention. However, event sparsity, noise, and limited texture still hinder robust local feature learning. Deploying such methods on resource-constrained platforms such as unmanned aerial vehicles also requires balancing accuracy and energy efficiency. To address these challenges, this paper proposes \textbf{E-S2Feat}, a spiking neural network framework for event-based local feature detection and description. The framework jointly optimizes local feature learning from the perspectives of feature representation and selection. First, a module-specific spiking activation mechanism preserves fine-grained structural cues and discriminative information under low-bit, energy-efficient inference, thereby improving overall representation fidelity. Furthermore, a semantic-guided feature modulation mechanism leverages semantic priors to refine keypoint response distributions and enhance local descriptor discriminability, thereby guiding the model to extract local features with greater geometric stability and stronger discriminative capability. Experiments on the ECD and EDS datasets show that the proposed method significantly outperforms baseline methods such as SuperEvent in pose estimation accuracy. It also achieves accuracy comparable to its artificial neural network counterpart while delivering an approximately 4.8-fold improvement in theoretical computational energy efficiency. Visual-inertial odometry experiments on the TUM-VIE dataset further verify the effectiveness and practical application potential of the proposed method in complete SLAM systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。