arXiv:2507.07903cs.CVeess.IV2025-07中稿 · the DSD 2025 confe…被引 2

在FPGA上实现低功耗视觉里程计,实时处理640x480图像达54帧/秒。

Hardware-Aware Feature Extraction Quantisation for Real-Time Visual Odometry on FPGA Platforms

  • 基于量化SuperPoint网络,实现无监督特征提取
  • 在Zynq UltraScale+ FPGA上达54帧/秒,优于现有方案
  • 适合资源受限的移动或嵌入式导航系统

精确位姿估计对自主平台(如地面车辆、水面船只和空中无人机)的现代导航系统至关重要。在此背景下,视觉同时定位与建图(VSLAM)依赖于从视觉输入中可靠提取显著特征点。本文提出一种嵌入式实现方案,采用量化版SuperPoint卷积神经网络,实现无监督特征检测与描述。目标是在保持高检测质量的前提下最小化计算开销,以支持在资源受限的移动或嵌入式系统上的高效部署。我们在AMD/Xilinx Zynq UltraScale+ FPGA系统级芯片平台上实现该方案,评估了深度学习处理单元(DPUs)性能,并使用Brevitas库与FINN框架完成模型量化与硬件感知优化。实验表明,可实现640×480像素图像最高54帧/秒的处理速度,优于当前最优方案。我们基于TUM数据集开展实验,分析不同量化技术对视觉里程计任务中模型精度与性能的影响。

原文摘要 · Abstract (English)

Accurate position estimation is essential for modern navigation systems deployed in autonomous platforms, including ground vehicles, marine vessels, and aerial drones. In this context, Visual Simultaneous Localisation and Mapping (VSLAM) - which includes Visual Odometry - relies heavily on the reliable extraction of salient feature points from the visual input data. In this work, we propose an embedded implementation of an unsupervised architecture capable of detecting and describing feature points. It is based on a quantised SuperPoint convolutional neural network. Our objective is to minimise the computational demands of the model while preserving high detection quality, thus facilitating efficient deployment on platforms with limited resources, such as mobile or embedded systems. We implemented the solution on an FPGA System-on-Chip (SoC) platform, specifically the AMD/Xilinx Zynq UltraScale+, where we evaluated the performance of Deep Learning Processing Units (DPUs) and we also used the Brevitas library and the FINN framework to perform model quantisation and hardware-aware optimisation. This allowed us to process 640 x 480 pixel images at up to 54 fps on an FPGA platform, outperforming state-of-the-art solutions in the field. We conducted experiments on the TUM dataset to demonstrate and discuss the impact of different quantisation techniques on the accuracy and performance of the model in a visual odometry task.

视觉里程计FPGA模型量化嵌入式部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。