arXiv:2601.09248cs.CVcs.AI2026-01

用事件相机+新型变分自编码器实现低功耗高鲁棒的室内定位

Hybrid guided variational autoencoder for visual place recognition

  • 结合事件相机与脉冲神经网络,适配低功耗神经形态硬件
  • 在16个场景中实现接近顶尖模型的识别准确率,光照变化下仍稳定
  • 对未知场景有强泛化能力,适合移动机器人在复杂环境导航

自动驾驶车辆、机器人和无人机需在包括无GPS信号的室内环境中精确定位。视觉位置识别(VPR)通过比对先前见过的图像来估计当前位置。现有先进VPR模型内存占用高,难以部署于移动端;而轻量模型则缺乏鲁棒性和泛化能力。本文提出一种基于事件视觉传感器的新型引导式变分自编码器(VAE),其编码器采用兼容低功耗、低延迟神经形态硬件的脉冲神经网络。该模型在新构建的室内VPR数据集上成功解耦出16个不同位置的视觉特征,分类性能达到当前顶尖水平,且在多种光照条件下表现稳健。在未知场景的新输入测试中,模型仍能有效区分不同位置,证明其具备学习位置本质特征的强泛化能力。该紧凑、鲁棒且具泛化能力的引导式VAE为移动机器人在已知与未知室内环境中的导航提供了极具前景的解决方案。

原文摘要 · Abstract (English)

Autonomous agents such as cars, robots and drones need to precisely localize themselves in diverse environments, including in GPS-denied indoor environments. One approach for precise localization is visual place recognition (VPR), which estimates the place of an image based on previously seen places. State-of-the-art VPR models require high amounts of memory, making them unwieldy for mobile deployment, while more compact models lack robustness and generalization capabilities. This work overcomes these limitations for robotics using a combination of event-based vision sensors and an event-based novel guided variational autoencoder (VAE). The encoder part of our model is based on a spiking neural network model which is compatible with power-efficient low latency neuromorphic hardware. The VAE successfully disentangles the visual features of 16 distinct places in our new indoor VPR dataset with a classification performance comparable to other state-of-the-art approaches while, showing robust performance also under various illumination conditions. When tested with novel visual inputs from unknown scenes, our model can distinguish between these places, which demonstrates a high generalization capability by learning the essential features of location. Our compact and robust guided VAE with generalization capabilities poses a promising model for visual place recognition that can significantly enhance mobile robot navigation in known and unknown indoor environments.

视觉定位事件相机变分自编码器机器人导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。