用自监督学习的隐式动作表示,大幅减少标注数据需求,提升相机位姿估计精度。
LA-Pose: Latent Action Pretraining Meets Pose Estimation

- 通过逆动力学模型学习视频中的隐式动作特征
- 在Waymo和PandaSet上比现有方法高10%以上准确率,仅需极少标注数据
- 适合数据稀缺场景下的实时位姿估计应用
本文从自监督预训练角度重新审视相机位姿估计,采用逆动力学与前向动力学模型学习潜在动作表示,类似大规模驾驶视频中的Genie方法。以往方法将潜在动作用于世界模型或策略网络中的动作条件,而本工作提出LA-Pose,将潜在动作特征作为相机位姿估计器的输入,并在少量高质量3D标注数据上微调。该方法实现高精度且可泛化的位姿预测,同时保持前馈效率。在多个驾驶基准测试中,LA-Pose在使用远少于现有方法的标注数据情况下,于Waymo和PandaSet上均取得超过10%的精度提升,性能优于当前先进方法。据我们所知,这是首个证明逆动力学自监督学习在位姿估计中有效性的研究。
原文摘要 · Abstract (English)
This paper revisits camera pose estimation through the lens of self-supervised pretraining, focusing on inverse-dynamics pretraining as a scalable alternative to the current trend of fully supervised training with 3D annotations. Concretely, we employ inverse- and forward-dynamics models to learn latent action representations, similar to Genie from large-scale driving videos. Our idea is simple yet effective. Existing methods use latent actions in their original capacity, that is, as action conditioning of world-models or as proxies of robot action parameters in policy networks. Our method, dubbed LA-Pose, repurposes the latent action features as inputs to a camera pose estimator, finetuned on a limited set of high-quality 3D annotations. This formulation enables accurate and generalizable pose prediction while maintaining feed-forward efficiency. Extensive experiments on driving benchmarks show that LA-Pose achieves competitive and even superior performance to state-of-the-art methods while using orders of magnitude less labeled data. Concretely, on the Waymo and PandaSet benchmarks, LA-Pose achieves over 10% higher pose accuracy than recent feed-forward methods. To our knowledge, this work is the first to demonstrate the power of inverse-dynamics self-supervised learning for pose estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。