grayscale比彩色更适合作为视觉定位的输入,尤其在环境变化大时表现更好
One Channel to Rule Them All: Rethinking Input Representation for Visual Place Recognition

- 用灰度图像替代彩色输入,简化模型并提升鲁棒性
- 灰度模型在多个基准上召回率平均达82.4%,优于彩色模型的81.2%
- 适合资源受限场景,节省存储与带宽,部署更高效
视觉定位(VPR)是机器人长期定位与建图的核心,现有系统普遍依赖彩色输入,隐含认为颜色对全局定位不可或缺。本文挑战这一假设,在真实世界外观变化下,考察色彩信息在不同训练方式、模型结构和标准基准中的作用。结果表明,灰度图像在多数情况下可媲美彩色表现,并在严重外观变化下更优,因颜色不变性难以充分学习;而颜色仅在存在持久且区分性强的色度线索时才带来增益。在选定基准上,全灰度训练的MixVPR模型平均召回率@1达82.4%,高于其彩色版本的81.2%。部分轻量级灰度模型参数减少60%,仍优于更重的彩色模型。灰度还具备存储、带宽和适配资源受限系统的实际优势。结论:在光照、天气、季节和场景变化大的全局VPR任务中,颜色贡献极小,灰度已足够实现可靠定位。
原文摘要 · Abstract (English)
Visual Place Recognition (VPR) is fundamental to long-term robot localization and SLAM, yet current systems overwhelmingly rely on RGB input, implicitly assuming color is necessary for global place recognition. We challenge this assumption, investigating the role of chromatic information across training regimes, model architectures and standard benchmarks under real-world appearance variation. We find that grayscale matches RGB performance generally and outperforms it under severe appearance shifts where color invariance is insufficiently learned, while color provides meaningful gains only where persistent and discriminative chromatic cues are present. Across selected benchmarks, a fully gray-trained MixVPR model achieves an average 82.4% Recall@1 compared to 81.2% for its RGB counterpart. In some cases, lightweight grayscale variants with 60% fewer parameters can outperform heavier RGB models. Grayscale further offers practical advantages in storage, bandwidth and alignment with resource-constrained systems. We conclude that for global VPR where scenes vary across illumination, weather, season and setting, color contributes minimally, and grayscale alone is sufficient for reliable place recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。