无需训练,用矩阵分解融合多条件特征,提升视觉定位精度
Joint Multi-Condition Representation Modelling via Matrix Factorisation for Visual Place Recognition
- 通过矩阵分解将多参考描述子统一为基础表示,实现投影残差匹配
- 在多光照多视角数据上,Recall@1提升约18%,未结构化数据增益达5%
- 轻量级设计适合实时部署,尤其适用于条件变化大的场景
针对多参考视觉定位(VPR)问题,本文提出一种无需训练、不依赖特定描述子的方法。通过矩阵分解将不同条件下的参考描述子联合建模为基表示,支持基于投影的残差匹配。同时构建了结构化基准SotonMV用于多视角VPR评估。在多光照、多视角数据上,方法相比单参考模型召回率提升约18%,优于现有多参考基线,在非结构化数据上仍实现约5%的增益,展现出强泛化能力且计算开销低。
原文摘要 · Abstract (English)
We address multi-reference visual place recognition (VPR), where reference sets captured under varying conditions are used to improve localisation performance. While deep learning with large-scale training improves robustness, increasing data diversity and model complexity incur extensive computational cost during training and deployment. Descriptor-level fusion via voting or aggregation avoids training, but often targets multi-sensor setups or relies on heuristics with limited gains under appearance and viewpoint change. We propose a training-free, descriptor-agnostic approach that jointly models places using multiple reference descriptors via matrix decomposition into basis representations, enabling projection-based residual matching. We also introduce SotonMV, a structured benchmark for multi-viewpoint VPR. On multi-appearance data, our method improves Recall@1 by up to ~18% over single-reference and outperforms multi-reference baselines across appearance and viewpoint changes, with gains of ~5% on unstructured data, demonstrating strong generalisation while remaining lightweight.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。