用单目相机实现无人机户外环境实时深度与语义分割。
Real-Time Monocular Scene Analysis for UAV in Outdoor Environments
- 提出Co-SemDepth模型,联合优化深度与语义估计。
- 在真实数据上验证,对海洋场景有良好泛化能力。
- 推荐扩散模型用于合成到真实图像风格迁移。
本论文利用安装在飞行机器人上的单目相机,在低空非结构化环境中预测深度图与语义图。提出一种联合深度学习架构Co-SemDepth,可准确快速完成两项任务,并在多个数据集上验证其有效性。神经网络训练需大量标注数据,而无人机领域此类数据稀缺。为此,本文构建了新合成数据集TopAir,包含不同高度俯视的室外图像,以填补数据空白。尽管合成数据训练便捷,但存在域迁移问题。通过系统分析发现,Co-SemDepth在深度估计上表现更优,TaskPrompter在语义分割上更佳;同时确定了提升泛化性能的训练数据组合。为缩小合成与真实域差距,探索了基于Cycle-GAN和扩散模型的图像风格迁移方法,结果显示扩散模型效果更优。最后聚焦海洋场景,使用自建合成海洋数据集MidSea训练Co-SemDepth,测试结果表明在SMD真实数据上表现良好,但在MIT数据集上仍需改进。
原文摘要 · Abstract (English)
In this thesis, we leverage monocular cameras on aerial robots to predict depth and semantic maps in low-altitude unstructured environments. We propose a joint deep-learning architecture, named Co-SemDepth, that can perform the two tasks accurately and rapidly, and validate its effectiveness on a variety of datasets. The training of neural networks requires an abundance of annotated data, and in the UAV field, the availability of such data is limited. We introduce a new synthetic dataset in this thesis, TopAir that contains images captured with a nadir view in outdoor environments at different altitudes, helping to fill the gap. While using synthetic data for the training is convenient, it raises issues when shifting to the real domain for testing. We conduct an extensive analytical study to assess the effect of several factors on the synthetic-to-real generalization. Co-SemDepth and TaskPrompter models are used for comparison in this study. The results reveal a superior generalization performance for Co-SemDepth in depth estimation and for TaskPrompter in semantic segmentation. Also, our analysis allows us to determine which training datasets lead to a better generalization. Moreover, to help attenuate the gap between the synthetic and real domains, image style transfer techniques are explored on aerial images to convert from the synthetic to the realistic style. Cycle-GAN and Diffusion models are employed. The results reveal that diffusion models are better in the synthetic to real style transfer. In the end, we focus on the marine domain and address its challenges. Co-SemDepth is trained on a collected synthetic marine data, called MidSea, and tested on both synthetic and real data. The results reveal good generalization performance of Co-SemDepth when tested on real data from the SMD dataset while further enhancement is needed on the MIT dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。