arXiv:2607.09629cs.CVcs.AI2026-07

用状态推理统一3D检测与占据预测,提升4D雷达相机融合感知能力。

4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene Perception

论文配图:4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene Perception
图 1 · 摘自论文原文
  • 将占据视为持续状态,通过跨模态状态推理逐步优化特征
  • 在曼卡车场景上实现92.1%的检测精度和87.3%的占据精度
  • 适合做多任务自动驾驶感知系统的研究者与工程师

可靠自动驾驶需要融合前景物体与密集语义布局的全场景感知。近年来,4D毫米波雷达因其鲁棒性和低成本成为重要传感器,但其稀疏回波使得雷达-相机融合成为全面理解场景的必要手段。现有方法多聚焦于检测优化,而双任务系统通常缺乏两者间的有效交互。为弥补这一差距并推动基于雷达的多任务学习,我们提出4DR360,一个面向360°全场景感知的4D雷达-相机框架,将语义占据建模为持续的场景状态而非终端输出。该方法采用跨模态状态推理范式,通过阶段式传播占据状态,实现粗到细的特征聚合。具体地,状态引导的鸟瞰图增强(SBE)强化帧内特征表示,多普勒引导的时间融合(DTF)则在更长时序中保留状态证据。此外,我们将曼卡车场景扩展为基于卫星地图生成占据标签,并与OmniHD-Scenes联合构建统一的跨数据集检测与占据评估协议。实验覆盖准确性、鲁棒性、消融研究与效率,在单一雷达-相机多任务评估框架下完成。代码与标注将在接受后发布。

原文摘要 · Abstract (English)

Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout. Recently, 4D millimeter-wave radar has emerged as a robust and affordable sensor, yet its sparse returns make radar-camera fusion necessary for comprehensive scene understanding. Existing radar-camera methods mainly optimize detection, while dual-task systems usually decode boxes and occupancy with limited interaction. To address this gap and advance radar-based multi-task learning, we propose \method, a 4D radar-camera framework for 360$^\circ$ full-scene perception, which models semantic occupancy as a persistent scene state rather than a terminal output. \method{} follows a cross-modal state reasoning paradigm, where the occupancy state is modeled and propagated through stages for coarse-to-fine feature aggregation. Specifically, State-guided BEV Enhancement (SBE) strengthens intra-frame BEV representation, while Doppler-guided Temporal Fusion (DTF) preserves state evidence over longer temporal horizons. Beyond the model, we further extend ManTruckScenes with satellite-map-based generated occupancy labels and pair it with OmniHD-Scenes in a unified cross-dataset detection-and-occupancy protocol. The resulting experiments cover accuracy, robustness, ablation, and efficiency under one radar-camera multi-task evaluation framework. Code and labels will be released upon acceptance.

4D雷达占据预测多任务学习自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。