融合语义与深度信息,提升自动驾驶中物体检测精度
Data Fusion of Semantic and Depth Information in the Context of Object Detection
- 用双目视觉生成视差图,结合Inception V2的Faster R-CNN进行3D定位
- 在自定义数据集上训练,实现对行人位置及距离的精确估计
- 适合研究多模态感知与自动驾驶感知系统优化的研究者
现代自动驾驶研究已取得显著进展。为确保安全,自动驾驶系统必须精准检测车辆周边物体。本文研究了物体(行人)的分类及其在自身3D坐标系下的位置估计,并测量了车辆与物体之间的距离。采用基于Inception V2的Faster Region-based Convolutional Neural Network(Faster R-CNN)进行物体分类。首先,在自定义数据集上训练网络以估计物体参考位置及与车辆的距离。从相机标定到距离计算,一系列先进的计算机视觉算法被应用于生成感兴趣区域的3D参考点。该过程的关键步骤是利用立体视觉原理生成视差图。
原文摘要 · Abstract (English)
Considerable study has already been conducted regarding autonomous driving in modern era. An autonomous driving system must be extremely good at detecting objects surrounding the car to ensure safety. In this paper, classification, and estimation of an object's (pedestrian) position (concerning an ego 3D coordinate system) are studied and the distance between the ego vehicle and the object in the context of autonomous driving is measured. To classify the object, faster Region-based Convolution Neural Network (R-CNN) with inception v2 is utilized. First, a network is trained with customized dataset to estimate the reference position of objects as well as the distance from the vehicle. From camera calibration to computing the distance, cutting-edge technologies of computer vision algorithms in a series of processes are applied to generate a 3D reference point of the region of interest. The foremost step in this process is generating a disparity map using the concept of stereo vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。