arXiv:2511.16521cs.CV2025-11被引 4

用移动设备一次扫描,同时完成室内建图和吊装摄像头校准。

YOWO: You Only Walk Once to Jointly Map An Indoor Scene and Register Ceiling-mounted Cameras

  • 让移动设备走一遍房间,同步采集自身与吊装摄像头数据。
  • 通过时间对齐轨迹,实现场景布局与摄像头位姿的联合优化。
  • 适用于需要精准定位的智能楼宇、无人配送等场景。

利用吊装摄像头(CMCs)进行室内视觉采集具有广泛应用前景,但将这些摄像头精确注册到场景布局中仍具挑战性。传统人工标定效率低且成本高,而基于视觉定位的自动方法在存在视觉模糊时效果不佳。为此,本文提出一种新方法,可联合完成室内场景建图与吊装摄像头注册。该方法通过配备头戴式RGB-D相机的移动代理遍历整个场景一次,同步触发吊装摄像头拍摄该移动代理。移动代理的视角视频生成世界坐标下的轨迹和场景布局,而吊装摄像头视频则提供伪尺度下的代理轨迹及相对位姿。通过时间戳对齐所有轨迹,即可将吊装摄像头相对位姿对齐至世界坐标场景布局。在此基础上,构建定制化因子图,实现自我相机位姿、场景布局与吊装摄像头位姿的联合优化。同时,我们构建了一个新数据集,为协同建图与吊装摄像头注册设定首个基准测试(https://sites.google.com/view/yowo/home)。实验表明,本方法不仅在统一框架内高效完成两项任务,还共同提升了性能表现,为下游位置感知应用提供了可靠工具。

原文摘要 · Abstract (English)

Using ceiling-mounted cameras (CMCs) for indoor visual capturing opens up a wide range of applications. However, registering CMCs to the target scene layout presents a challenging task. While manual registration with specialized tools is inefficient and costly, automatic registration with visual localization may yield poor results when visual ambiguity exists. To alleviate these issues, we propose a novel solution for jointly mapping an indoor scene and registering CMCs to the scene layout. Our approach involves equipping a mobile agent with a head-mounted RGB-D camera to traverse the entire scene once and synchronize CMCs to capture this mobile agent. The egocentric videos generate world-coordinate agent trajectories and the scene layout, while the videos of CMCs provide pseudo-scale agent trajectories and CMC relative poses. By correlating all the trajectories with their corresponding timestamps, the CMC relative poses can be aligned to the world-coordinate scene layout. Based on this initialization, a factor graph is customized to enable the joint optimization of ego-camera poses, scene layout, and CMC poses. We also develop a new dataset, setting the first benchmark for collaborative scene mapping and CMC registration (https://sites.google.com/view/yowo/home). Experimental results indicate that our method not only effectively accomplishes two tasks within a unified framework, but also jointly enhances their performance. We thus provide a reliable tool to facilitate downstream position-aware applications.

三维重建摄像头校准多源融合智能建筑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。