仅用单张图像,轻量模型实时估算3D物体尺寸与朝向。
MonoLite3D: Lightweight 3D Object Properties Estimation
- 基于单目图像的轻量化网络,适配资源受限设备。
- KITTI数据集上朝向估计准确率82.27%(中等难度)。
- 兼顾精度与速度,适合车载嵌入式部署。
可靠的环境感知对自动驾驶车辆至关重要。感知系统需在限定时间内获取周围物体的完整3D信息,包括尺寸、空间位置与方向。深度学习在感知系统中广泛应用,可将摄像头捕获的图像特征转化为有意义的语义信息。本文提出MonoLite3D网络,一种面向嵌入式设备的轻量级深度学习方法,专为计算资源有限的硬件环境设计。该方法仅通过单目图像即可估计3D物体的多维度属性,如尺寸与空间朝向。实验结果表明,在KITTI数据集的方向估计基准测试中,该方法在中等难度类别上达到82.27%的准确率,在困难类别上达到69.81%,同时满足实时性要求。
原文摘要 · Abstract (English)
Reliable perception of the environment plays a crucial role in enabling efficient self-driving vehicles. Therefore, the perception system necessitates the acquisition of comprehensive 3D data regarding the surrounding objects within a specific time constrain, including their dimensions, spatial location and orientation. Deep learning has gained significant popularity in perception systems, enabling the conversion of image features captured by a camera into meaningful semantic information. This research paper introduces MonoLite3D network, an embedded-device friendly lightweight deep learning methodology designed for hardware environments with limited resources. MonoLite3D network is a cutting-edge technique that focuses on estimating multiple properties of 3D objects, encompassing their dimensions and spatial orientation, solely from monocular images. This approach is specifically designed to meet the requirements of resource-constrained environments, making it highly suitable for deployment on devices with limited computational capabilities. The experimental results validate the accuracy and efficiency of the proposed approach on the orientation benchmark of the KITTI dataset. It achieves an impressive score of 82.27% on the moderate class and 69.81% on the hard class, while still meeting the real-time requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。