用双目视觉无源融合图像与深度,精准估算非合作航天器6自由度姿态。
Cross-Modal RGB-D Fusion Transformer for 6D Pose Estimation of Non-Cooperative Spacecraft with Stereo-Derived Depth
- 设计TSCA-Stereo网络处理太空图像弱纹理和强光照问题。
- 融合RGB与深度特征的Transformer使姿态估计误差达0.0419米、0.8632度。
- 适用于光照恶劣、资源受限的在轨服务场景,适合航天器自主导航应用。
在轨服务与主动清除空间碎片需对非合作航天器进行可靠姿态估计,以提供自主视觉导航所需的位置与姿态信息。基于学习的单目方法虽广泛应用,但存在固有的深度模糊问题,在轨道常见严苛光照下易失效。主动深度传感器虽可解决几何歧义,但功耗与质量过大,不适用于多数航天平台。本文提出一种被动双目视觉框架,实现非合作航天器六自由度(6-DOF)姿态估计。构建了名为TSCA-Stereo的双目匹配网络,以应对太空图像中常见的弱纹理、镜面反射和剧烈光照变化。引入跨模态融合Transformer,自适应结合RGB外观信息与立体深度特征,支持鲁棒姿态恢复。同时构建了一个合成双目多模态数据集,包含不同光照条件、姿态配置和噪声水平下的立体视差图与6-DOF姿态标注。实验表明,TSCA-Stereo在所有评估指标上均优于基线。完整姿态估计流程在多种成像条件下实现平均平移误差0.0419米、平均旋转误差0.8632°,验证了被动双目方案在复杂太空视觉环境中的有效性与鲁棒性。
原文摘要 · Abstract (English)
On-orbit servicing and active debris removal involving non-cooperative spacecraft require reliable pose estimation to supply accurate position and orientation data for autonomous visual navigation. Learning-based monocular methods have seen widespread adoption in spacecraft pose estimation, yet they suffer from an intrinsic depth ambiguity problem and tend to fail under the harsh illumination conditions routinely encountered in orbit. Active depth sensors could in principle address the geometric ambiguity, but their power and mass requirements make them poorly suited to most spacecraft platforms. This work addresses these issues through a passive stereo vision framework for six-degree-of-freedom (6-DOF) pose estimation of non-cooperative spacecraft. A binocular stereo matching network called TSCA-Stereo is developed to cope with weak-texture surfaces, specular highlights, and severe lighting variations typical of space imagery. A cross-modal fusion Transformer is introduced to combine RGB appearance information with stereo depth features in an adaptive manner, supporting reliable pose recovery. A synthetic binocular multimodal dataset is also built for the experiments, covering stereo disparity maps and 6-DOF pose annotations across a range of illumination scenarios, attitude configurations, and noise levels. Experimental results show that TSCA-Stereo outperforms the baseline across every evaluated metric on this space-specific dataset. The full pose estimation pipeline achieves a mean translation error of 0.0419 m and a mean orientation error of 0.8632° under varied imaging conditions, confirming that the passive stereo approach is both effective and resilient when operating under the demanding visual conditions of the space environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。