用极稀疏深度数据完成未知野外环境的稠密深度估计
Depth Completion in Unseen Field Robotics Environments Using Extremely Sparse Depth Measurements
- 基于合成数据训练,利用稀疏深度输入预测稠密度量深度
- 在真实野外场景中表现优异,端到端延迟仅53毫秒
- 适合嵌入式平台部署,适用于无纹理或复杂光照的野外机器人
在非结构化环境中运行的自主场外机器人需要可靠的感知能力以确保安全与稳定作业。尽管单目深度估计的进展展示了低成本摄像头作为深度传感器的潜力,但其在场外机器人中的应用仍受限于缺乏可靠尺度信息、纹理不足或低纹理条件,以及大规模数据集的匮乏。为此,我们提出一种深度补全模型,该模型在合成数据上训练,并利用深度传感器提供的极稀疏测量值,在未见过的场外机器人环境中预测稠密的度量深度。针对场外机器人定制的合成数据生成流程可创建多个真实感强的数据集用于训练,该方法结合运动结构(Structure from Motion)生成的带纹理3D网格与基于新视角合成的逼真渲染技术,模拟多样化的场外机器人场景。我们的方法在Nvidia Jetson AGX Orin平台上实现每帧53毫秒的端到端延迟,支持嵌入式平台实时部署。大量实验证明其在多种真实场外机器人场景中具备竞争力的表现。
原文摘要 · Abstract (English)
Autonomous field robots operating in unstructured environments require robust perception to ensure safe and reliable operations. Recent advances in monocular depth estimation have demonstrated the potential of low-cost cameras as depth sensors; however, their adoption in field robotics remains limited due to the absence of reliable scale cues, ambiguous or low-texture conditions, and the scarcity of large-scale datasets. To address these challenges, we propose a depth completion model that trains on synthetic data and uses extremely sparse measurements from depth sensors to predict dense metric depth in unseen field robotics environments. A synthetic dataset generation pipeline tailored to field robotics enables the creation of multiple realistic datasets for training purposes. This dataset generation approach utilizes textured 3D meshes from Structure from Motion and photorealistic rendering with novel viewpoint synthesis to simulate diverse field robotics scenarios. Our approach achieves an end-to-end latency of 53 ms per frame on a Nvidia Jetson AGX Orin, enabling real-time deployment on embedded platforms. Extensive evaluation demonstrates competitive performance across diverse real-world field robotics scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。