无需基础设施,用深度学习实现飞机外检相机位姿估计与图像定位。
CNN-Based Camera Pose Estimation and Localisation of Scan Images for Aircraft Visual Inspection
- 用合成数据微调CNN,自预测相机自身位姿。
- 实测位姿误差低于0.24米和2度,适用于户外快速检测。
- 适合机场现场部署,不依赖无人机或接触飞机表面。
通用视觉检查是定期用于检测和定位商用飞机外部明显损伤的手动流程。为减少飞机停机时间,需在登机口完成此过程,自动化可降低对人力的依赖。但现有定位方法多需基础设施,难以在无控制的室外环境及约2小时的机坪周转时间内实施。此外,多数航空公司禁止接触飞机表面或在航班间使用无人机,且限制进入商业飞机。为此,本文提出一种无需基础设施、易于部署的现场方法,用于估算全景-俯仰-变焦相机的位姿并定位扫描图像。该方法利用执行检测任务的同一相机,通过仅在合成图像上微调的深度卷积神经网络预测自身位姿。采用领域随机化生成训练数据,并结合飞机几何信息优化损失函数以提升精度。同时提出从初始化到扫描路径规划再到图像精确定位的完整工作流程。通过真实飞机实验验证,所有场景下均实现小于0.24米和2度的均方根位姿估计误差。
原文摘要 · Abstract (English)
General Visual Inspection is a manual inspection process regularly used to detect and localise obvious damage on the exterior of commercial aircraft. There has been increasing demand to perform this process at the boarding gate to minimise the downtime of the aircraft and automating this process is desired to reduce the reliance on human labour. Automating this typically requires estimating a camera's pose with respect to the aircraft for initialisation but most existing localisation methods require infrastructure, which is very challenging in uncontrolled outdoor environments and within the limited turnover time (approximately 2 hours) on an airport tarmac. Additionally, many airlines and airports do not allow contact with the aircraft's surface or using UAVs for inspection between flights, and restrict access to commercial aircraft. Hence, this paper proposes an on-site method that is infrastructure-free and easy to deploy for estimating a pan-tilt-zoom camera's pose and localising scan images. This method initialises using the same pan-tilt-zoom camera used for the inspection task by utilising a Deep Convolutional Neural Network fine-tuned on only synthetic images to predict its own pose. We apply domain randomisation to generate the dataset for fine-tuning the network and modify its loss function by leveraging aircraft geometry to improve accuracy. We also propose a workflow for initialisation, scan path planning, and precise localisation of images captured from a pan-tilt-zoom camera. We evaluate and demonstrate our approach through experiments with real aircraft, achieving root-mean-square camera pose estimation errors of less than 0.24 m and 2 degrees for all real scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。