arXiv:2607.26817cs.ROcs.CV2026-07

不依赖射线匹配,通过粗到精框架实现高精度室内定位

From Uncertainty to Determinism: Coarse-to-Fine Visual Floorplan Localization without Ray Matching

论文配图:From Uncertainty to Determinism: Coarse-to-Fine Visual Floorplan Localization without Ray Matching
图 1 · 摘自论文原文
  • 用图像条件的扩散模型建模多模态位姿分布,从不确定性中推导候选位置
  • 在候选位置附近预测亚米级位姿残差,消除结构歧义,定位精度达亚米级
  • 无需离线地图预处理或测试时查表,适合真实场景部署

视觉楼层平面定位(FLoc)通过将第一人称图像与简约结构地图匹配,成为一种有前景的室内定位方案。然而,由于跨模态信息不对称和重复的室内布局,视觉FLoc面临多模态位姿分布的根本挑战——视觉上相同的观测可能对应空间分离的不同位置。现有基于射线匹配的方法通过显式预测稀疏几何或语义射线来解决此问题,但会引入信息损失,并需资源密集型预处理及推理时的穷举匹配。本文跳过中间射线匹配范式,提出一种从粗到精的视觉FLoc框架,逐步从不确定性走向确定性。在粗粒度阶段,设计图像条件的位姿扩散模型,参数化连续的多模态位姿分布,将随机初始化的位姿粒子引导至不同候选模式。在精炼阶段,提出局部精修器,从候选中心的楼层平面图像块中预测受限的亚米级位姿残差,此时结构歧义已大幅消除。该方法在S3D(完整版)和ZInD基准上均达到当前最优精度与鲁棒性,无需任何离线地图预处理或测试时查找表。

原文摘要 · Abstract (English)

Visual Floorplan Localization (FLoc) has emerged as a promising solution for indoor localization by matching egocentric images against minimalist structural maps. However, due to cross-modal information asymmetry and repetitive indoor layouts, visual FLoc is fundamentally challenged by multimodal pose distributions, where visually identical observations map to distinct, spatially separated locations. Existing ray-matching-based methods tackle this by explicitly predicting sparse geometric or semantic rays, which inherently incur information loss and demand resource-intensive preprocessing alongside exhaustive matching during inference. In this paper, we bypass the intermediate ray-matching paradigm and propose a coarse-to-fine visual FLoc framework that progresses from uncertainty to determinism. In the coarse stage, we design an image-conditioned pose diffusion model to parameterize the continuous multimodal pose distribution, effectively routing stochastically initialized pose particles toward distinct candidate modes. In the refinement stage, we propose a localized refiner that predicts bounded sub-meter pose residuals from candidate-centered floorplan crops, where structural ambiguities are largely eliminated. Our method effectively balances global multi-hypothesis tracking and local sub-meter refinement without requiring any offline map preprocessing or test-time lookup tables. Comprehensive results on the S3D (full) and ZInD benchmarks demonstrate that our approach achieves state-of-the-art accuracy and robustness.

室内定位扩散模型多模态无预处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。