arXiv:2504.05649cs.CV2025-04被引 1

用单帧调频连续波激光雷达预测物体未来位置,提升自动驾驶响应速度。

POD: Predictive Object Detection with Single-Frame FMCW LiDAR Point Cloud

  • 基于激光雷达径向速度生成虚拟未来点,构建双帧点云
  • 在单帧数据下实现未来3秒内物体位置与尺寸的准确预测
  • 适合对实时性要求高的自动驾驶感知系统

基于激光雷达的3D目标检测是自动驾驶中的基础任务。本文探索了调频连续波(FMCW)激光雷达在自主感知中的独特优势。给定带有径向速度测量的单帧FMCW点云,我们期望探测器仅凭当前帧传感器数据即可预测物体短时未来的位姿,并快速响应潜在危险。为此,我们将标准目标检测任务拓展为新型任务——预测性目标检测(POD),旨在仅依据当前观测预测物体短时未来的位置与尺寸。传统运动预测需历史传感器信息以建模时序上下文,而本方法避免多帧历史信息,显著提升对突发危险的响应速度。FMCW激光雷达的核心优势在于每个反射点均携带径向速度。我们提出一种新型POD框架:通过射线投射机制生成虚拟未来点,构建包含当前帧与虚拟未来帧的双帧点云,利用稀疏4D编码器编码该两帧体素特征。随后按时间索引分离4D体素特征,重映射为两个鸟瞰图(BEV)特征:一个解码用于标准当前帧目标检测,另一个用于未来预测检测。在自建数据集上的大量实验表明,所提POD框架在标准检测与预测检测性能上均达到当前最优水平。

原文摘要 · Abstract (English)

LiDAR-based 3D object detection is a fundamental task in the field of autonomous driving. This paper explores the unique advantage of Frequency Modulated Continuous Wave (FMCW) LiDAR in autonomous perception. Given a single frame FMCW point cloud with radial velocity measurements, we expect that our object detector can detect the short-term future locations of objects using only the current frame sensor data and demonstrate a fast ability to respond to intermediate danger. To achieve this, we extend the standard object detection task to a novel task named predictive object detection (POD), which aims to predict the short-term future location and dimensions of objects based solely on current observations. Typically, a motion prediction task requires historical sensor information to process the temporal contexts of each object, while our detector's avoidance of multi-frame historical information enables a much faster response time to potential dangers. The core advantage of FMCW LiDAR lies in the radial velocity associated with every reflected point. We propose a novel POD framework, the core idea of which is to generate a virtual future point using a ray casting mechanism, create virtual two-frame point clouds with the current and virtual future frames, and encode these two-frame voxel features with a sparse 4D encoder. Subsequently, the 4D voxel features are separated by temporal indices and remapped into two Bird's Eye View (BEV) features: one decoded for standard current frame object detection and the other for future predictive object detection. Extensive experiments on our in-house dataset demonstrate the state-of-the-art standard and predictive detection performance of the proposed POD framework.

激光雷达目标检测预测感知自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。