用建筑平面图的结构信息提升机器人定位精度
COMPASS: COmpact Multi-channel Prior-map And Scene Signature for Floor-Plan-Based Visual Localization

- 设计多通道径向描述符融合几何与语义信息
- 跨模态匹配验证:图像与平面图窗口墙模式高度一致
- 适合做视觉定位的科研人员和工程师参考
建筑平面图包含丰富的环境几何与语义信息,但现有定位方法大多忽略这些信息。为此,我们提出COMPASS算法,利用双鱼眼相机与平面图中的几何和语义先验进行位姿估计。受激光雷达场景上下文描述符启发,设计多通道径向描述符,编码位置周围的布局:从平面图出发,在360个方位角方向投射射线,生成五个通道——归一化距离、结构命中类型(墙、窗或开口)、距离梯度、逆距离及局部距离方差。从图像侧,通过检测鱼眼影像中的结构元素填充相同描述符结构。提出一种窗口检测算法,利用线段检测器结合垂直边缘聚类与亮度验证识别窗框,并通过鱼眼相机模型投影至方位角,生成视觉描述符的命中类型通道。作为概念验证,基于Hilti-Trimble SLAM Challenge 2026数据集,在单一定位点生成描述符,结果显示两摄像头首帧提取的墙窗模式与平面图描述符高度匹配,证实跨模态结构匹配的可行性。
原文摘要 · Abstract (English)
Architectural floor plans are widely available priors which contain not only geometry but also the semantic information of the environment, yet existing localization methods largely ignore this semantic information. To address this, we present COMPASS, an algorithm that exploits both geometric and semantic priors from floor plans to estimate the pose of a robot equipped with dual fisheye cameras. Inspired by scan context descriptor from LiDAR-based place recognition, we design a multi-channel radial descriptor that encodes the geometric layout surrounding a position. From the floor plan, rays are cast in 360 azimuth bins and the results are encoded into five channels: normalized range, structural hit type (wall, window, or opening), range gradient, inverse range, and local range variance. From the image side, the same descriptor structure is populated by detecting structural elements in the fisheye imagery. As a first step toward full cross-modal matching, we present a window detection algorithm for fisheye images that uses a line segment detector to identify window frames via vertical edge clustering and brightness verification. Detected windows are projected to azimuthal bearings through the fisheye camera model, producing the hit-type channel of the visual descriptor. As a proof of concept, we generate both descriptors at a single known pose from the Hilti-Trimble SLAM Challenge 2026 dataset and demonstrate that the wall-window pattern extracted from the first frame of each camera closely matches the floor plan descriptor, validating the feasibility of cross-modal structural matching.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。