arXiv:2509.06000cs.CV2025-09被引 2

用视觉变压器+运动信息提升单目航天器位姿估计精度

Motion Aware ViT-based Framework for Monocular 6-DoF Spacecraft Pose Estimation

  • 引入运动感知热图与光流,捕捉航天任务中的动态信息
  • 在SPADES-RGB上2D关键点定位误差降低12.3%,6-DoF位姿估计更准确
  • 跨数据分布测试表现良好,适合真实空间任务部署

单目6自由度位姿估计在多个航天任务中至关重要。现有方法多依赖静态图像中的关键点定位,未能利用航天操作中固有的时间信息。本文将人体姿态估计中的深度学习框架迁移至航天器位姿估计领域,结合运动感知热图与预训练光流模型,捕获运动动态。通过视觉变压器(ViT)编码器提取图像特征,并融合光流提供的运动线索以定位2D关键点,再基于已知的2D-3D对应关系,使用透视n点(PnP)求解器恢复6-DoF位姿。我们在SPADES-RGB数据集上训练并评估该方法,进一步在SPARK-2024的真实与合成数据上测试其泛化能力。结果表明,相比单图像基线,本方法在2D关键点定位和6-DoF位姿估计上均取得显著提升,且在不同数据分布下展现出良好的泛化性能。

原文摘要 · Abstract (English)

Monocular 6-DoF pose estimation plays an important role in multiple spacecraft missions. Most existing pose estimation approaches rely on single images with static keypoint localisation, failing to exploit valuable temporal information inherent to space operations. In this work, we adapt a deep learning framework from human pose estimation to the spacecraft pose estimation domain that integrates motion-aware heatmaps and optical flow to capture motion dynamics. Our approach combines image features from a Vision Transformer (ViT) encoder with motion cues from a pre-trained optical flow model to localise 2D keypoints. Using the estimates, a Perspective-n-Point (PnP) solver recovers 6-DoF poses from known 2D-3D correspondences. We train and evaluate our method on the SPADES-RGB dataset and further assess its generalisation on real and synthetic data from the SPARK-2024 dataset. Overall, our approach demonstrates improved performance over single-image baselines in both 2D keypoint localisation and 6-DoF pose estimation. Furthermore, it shows promising generalisation capabilities when testing on different data distributions.

位姿估计视觉变压器运动建模航天器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。