arXiv:2606.22834cs.CV2026-06

用几何关系指导相机精准拍摄平面区域,仅需少量标注即可实现高精度定位。

Homographic Navigation: Geometry-Driven Camera Guidance for Deterministic Planar Capture

论文配图:Homographic Navigation: Geometry-Driven Camera Guidance for Deterministic Planar Capture
图 1 · 摘自论文原文
  • 以单张参考图生成无限合成数据,通过稀疏关键点预测联合识别与定位物体
  • 两阶段推理+稳定变形训练,提升低分辨率输入下的定位精度,尤其在高精度场景
  • 输出关键点置信度,适合需要可靠几何引导的机器人或视觉系统应用

我们提出同构导航(Homographic Navigation),一种以几何为核心驱动相机采集平面区域的框架。不同于将单应性作为输出,我们将其作为组织变量,统一学习、对齐与评估流程。仅需一张标注参考图像,即可通过同构增强生成无限合成训练数据,并训练一个单次推理模型,通过稀疏关键点预测实现对多个带矩形平面目标的实物(artifacts)的联合识别与定位。为应对有限输入分辨率下的精度挑战,提出两阶段推理机制:先全局检测再局部精修,并引入稳定变形(Stable Warp)训练策略,显著提升高精度场景下的准确性。模型还为每个关键点及样本整体预测置信度。实验表明,仅需极少监督即可实现精确平面对齐,为几何驱动的相机引导及未来从真实视频中学习奠定基础。

原文摘要 · Abstract (English)

We present homographic navigation, a geometry-centric framework for guiding camera acquisition toward precise capture of planar regions. Rather than treating homography as an output, we use it as an organizing variable that unifies learning, alignment, and evaluation. From a single annotated reference image, we generate unlimited synthetic training data via homographic augmentation and train a single-shot model for joint recognition and localization of multiple artifacts (physical objects with a rectangular planar target) through sparse keypoint prediction. To address precision under limited model input resolution, we introduce a two-pass inference scheme with global detection followed by localized refinement, and a Stable Warp training strategy that significantly improves accuracy, particularly in the high-precision regime. The model also predicts confidence estimates per predicted keypoint and per the whole sample. Experimental results demonstrate that accurate planar alignment can be achieved from minimal supervision, providing a foundation for geometry-driven camera guidance and future learning from in-the-wild video data.

几何引导相机控制关键点预测高精度定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。