用双曲空间建模全景图层次结构,提升视角到全景图的识别效率
HypeVPR: Exploring Hyperbolic Space for Perspective to Equirectangular Visual Place Recognition
- 在双曲空间构建分层特征聚合机制,捕捉全景与视角间的层级关系
- 相比传统方法,检索速度提升3倍以上,数据库存储减少60%
- 无需额外训练即可灵活调节精度与效率,适合实时导航场景
视觉环境具有天然的层次结构:全景图自然包含并组织多个视角图像。准确捕捉这种层次结构对视角到全景图(P2E)的视觉位置识别至关重要。本文提出HypeVPR,一种基于双曲空间的分层嵌入框架,专门解决P2E匹配挑战。该方法利用双曲空间表示层次结构的内在优势,使全景描述子同时编码全局上下文和局部细节。我们设计了分层特征聚合机制,在双曲空间中组织从局部到全局的特征表示。此外,HypeVPR的分层结构可自然实现精度-效率权衡,无需额外训练,且在不同图像类型间保持鲁棒匹配。该方法在保持竞争力性能的同时,显著加速检索并减少数据库存储需求。项目主页:https://suhan-woo.github.io/HypeVPR/
原文摘要 · Abstract (English)
Visual environments are inherently hierarchical, as a panoramic view naturally encompasses and organizes multiple perspective views within its field. Capturing this hierarchy is crucial for effective perspective-to-equirectangular (P2E) visual place recognition. In this work, we introduce HypeVPR, a hierarchical embedding framework in hyperbolic space specifically designed to address the challenges of P2E matching. HypeVPR leverages the intrinsic ability of hyperbolic space to represent hierarchical structures, allowing panoramic descriptors to encode both broad contextual information and fine-grained local details. To this end, we propose a hierarchical feature aggregation mechanism that organizes local-to-global feature representations within hyperbolic space. Furthermore, HypeVPR's hierarchical organization naturally enables flexible control over the accuracy-efficiency trade-off without additional training, while maintaining robust matching across different image types. This approach enables HypeVPR to achieve competitive performance while significantly accelerating retrieval and reducing database storage requirements. Project page: https://suhan-woo.github.io/HypeVPR/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。