arXiv:2509.04948cs.ROcs.CV2025-09

仅用单张图像实现机器人在办公室环境的精准定位

Towards an Accurate and Effective Robot Vision (The Problem of Topological Localization for Mobile Robots)

  • 基于视角彩色相机图像,不依赖图像序列连续性
  • 验证多种视觉描述子,发现配置组合决定定位效果
  • 适用于光照变化大、路线长的真实场景定位任务

拓扑定位是移动机器人完成任务的基础,但受感知模糊、传感器噪声和光照变化影响,视觉定位与场景识别极具挑战。本文在办公室环境中,仅使用安装于机器人平台上的透视彩色相机采集的图像,不依赖图像序列的时间连续性,研究拓扑定位问题。系统评估了包括颜色直方图、SIFT、ASIFT、RGB-SIFT及受文本检索启发的词袋视觉模型在内的多种先进视觉描述子,通过标准评价指标与可视化手段,对特征、相似度度量和分类器进行了系统性定量比较。结果表明,外观描述子、相似度度量与分类器的合理配置显著提升定位性能。该配置在ImageCLEF视觉任务中进一步验证,成功识别新图像序列的最可能位置。未来工作将探索分层模型、排序方法与特征融合,以构建更鲁棒的定位系统,在降低训练与运行开销的同时避免维度灾难,最终实现跨光照变化、长距离的集成式实时定位。

原文摘要 · Abstract (English)

Topological localization is a fundamental problem in mobile robotics, since robots must be able to determine their position in order to accomplish tasks. Visual localization and place recognition are challenging due to perceptual ambiguity, sensor noise, and illumination variations. This work addresses topological localization in an office environment using only images acquired with a perspective color camera mounted on a robot platform, without relying on temporal continuity of image sequences. We evaluate state-of-the-art visual descriptors, including Color Histograms, SIFT, ASIFT, RGB-SIFT, and Bag-of-Visual-Words approaches inspired by text retrieval. Our contributions include a systematic, quantitative comparison of these features, distance measures, and classifiers. Performance was analyzed using standard evaluation metrics and visualizations, extending previous experiments. Results demonstrate the advantages of proper configurations of appearance descriptors, similarity measures, and classifiers. The quality of these configurations was further validated in the Robot Vision task of the ImageCLEF evaluation campaign, where the system identified the most likely location of novel image sequences. Future work will explore hierarchical models, ranking methods, and feature combinations to build more robust localization systems, reducing training and runtime while avoiding the curse of dimensionality. Ultimately, this aims toward integrated, real-time localization across varied illumination and longer routes.

机器人视觉图像定位特征描述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。