arXiv:2505.16447cs.CV2025-05

轻量化视觉定位模型,运行时可动态降算力40%且不丢精度。

TAT-VPR: Ternary Adaptive Transformer for Dynamic and Efficient Visual Place Recognition

  • 用三值权重+可学习稀疏门控,动态调节计算量
  • 运行时降低40%计算量,Recall@1性能不变
  • 适合微小型无人机和嵌入式设备部署

TAT-VPR 是一种三值量化变压器,为视觉 SLAM 回环检测带来动态的精度-效率权衡。通过融合三值权重与可学习激活稀疏门控,模型可在运行时将计算量减少高达40%,同时保持性能(Recall@1)不下降。所提出的两阶段知识蒸馏流程有效保留了描述子质量,使模型能够在微小型无人机及嵌入式 SLAM 系统上运行,并达到当前最优的定位精度。

原文摘要 · Abstract (English)

TAT-VPR is a ternary-quantized transformer that brings dynamic accuracy-efficiency trade-offs to visual SLAM loop-closure. By fusing ternary weights with a learned activation-sparsity gate, the model can control computation by up to 40% at run-time without degrading performance (Recall@1). The proposed two-stage distillation pipeline preserves descriptor quality, letting it run on micro-UAV and embedded SLAM stacks while matching state-of-the-art localization accuracy.

视觉定位轻量化Transformer嵌入式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。