利用深度与法向协同提升无人机单目安全着陆感知能力
VisLanding: Monocular 3D Perception for UAV Safe Landing via Depth-Normal Synergy
- 通过深度与法向联合优化,将着陆区识别转为二值语义分割
- 在跨域测试中准确率显著提升,且保持零样本泛化能力
- 可估算着陆区面积,适用于复杂未知环境下的实际飞行
本文提出VisLanding,一种基于单目3D感知的无人机安全着陆框架。针对复杂未知环境下无人机自主着陆的核心挑战,创新性地利用Metric3D V2模型的深度-法向协同预测能力,构建端到端安全着陆区(SLZ)估计框架。通过引入安全区分割分支,将着陆区估计任务转化为二值语义分割问题。模型基于无人机视角的WildUAV数据集进行微调与标注,并构建跨域评估数据集以验证鲁棒性。实验表明,VisLanding通过深度-法向联合优化机制显著提升了安全区识别精度,同时保留了Metric3D V2的零样本泛化优势。所提方法在跨域测试中表现出更优的泛化性与鲁棒性,并可通过融合预测的深度与法向信息估算着陆区面积,为实际应用提供关键决策支持。
原文摘要 · Abstract (English)
This paper presents VisLanding, a monocular 3D perception-based framework for safe UAV (Unmanned Aerial Vehicle) landing. Addressing the core challenge of autonomous UAV landing in complex and unknown environments, this study innovatively leverages the depth-normal synergy prediction capabilities of the Metric3D V2 model to construct an end-to-end safe landing zones (SLZ) estimation framework. By introducing a safe zone segmentation branch, we transform the landing zone estimation task into a binary semantic segmentation problem. The model is fine-tuned and annotated using the WildUAV dataset from a UAV perspective, while a cross-domain evaluation dataset is constructed to validate the model's robustness. Experimental results demonstrate that VisLanding significantly enhances the accuracy of safe zone identification through a depth-normal joint optimization mechanism, while retaining the zero-shot generalization advantages of Metric3D V2. The proposed method exhibits superior generalization and robustness in cross-domain testing compared to other approaches. Furthermore, it enables the estimation of landing zone area by integrating predicted depth and normal information, providing critical decision-making support for practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。