用条件变分自编码器实现图像驱动的多假设位姿估计
Conditional Variational Autoencoders for Probabilistic Pose Regression
- 基于条件变分自编码器构建图像到位姿的生成模型
- 可从后验分布采样,支持多个可能位姿输出
- 在重复结构环境下定位精度优于现有方法
机器人在失去跟踪时依赖视觉重定位来估计自身位姿。环境中重复结构带来的歧义是视觉重定位的主要挑战之一,这需要支持多假设的概率化方法。本文提出一种概率化方法,可根据观测图像预测相机位姿的后验分布。所提出的训练策略构建了一个以图像为条件的位姿生成模型,能够从位姿后验分布中采样。该方法理论基础扎实、结构简洁,在存在歧义的定位任务中表现优于现有方法。
原文摘要 · Abstract (English)
Robots rely on visual relocalization to estimate their pose from camera images when they lose track. One of the challenges in visual relocalization is repetitive structures in the operation environment of the robot. This calls for probabilistic methods that support multiple hypotheses for robot's pose. We propose such a probabilistic method to predict the posterior distribution of camera poses given an observed image. Our proposed training strategy results in a generative model of camera poses given an image, which can be used to draw samples from the pose posterior distribution. Our method is streamlined and well-founded in theory and outperforms existing methods on localization in presence of ambiguities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。