arXiv:2602.08006cs.CVcs.AI2026-02被引 4

首个直接从图像预测未来语义占位的框架,提升自动驾驶环境预判能力。

ForecastOcc: Vision-based Semantic Occupancy Forecasting

  • 联合预测未来多时序的语义占位,直接从图像学习时空特征。
  • 在Occ3D-nuScenes和SemanticKITTI上均超越基线,实现多视角与单目建模。
  • 创新设计时序跨注意力与2D-3D转换模块,适合自动驾驶场景理解研究者。

自动驾驶需要对未来环境状态进行几何与语义双重预测。现有基于视觉的占位预测方法主要关注静态与动态物体等运动类别,而语义信息普遍缺失。近期语义占位预测方法虽弥补此不足,但依赖独立网络生成的历史占位预测,导致误差累积,并阻碍从图像中直接学习时空特征。本文提出ForecastOcc,首个基于视觉的语义占位预测框架,可直接从过往相机图像中联合预测未来多个时间步的占位状态与语义类别,无需外部地图。我们在两个互补设置下评估:基于Occ3D-nuScenes的多视角预测,以及基于SemanticKITTI的单目预测,并建立该任务首个基准。通过适配两种2D预测模块构建首个基线。核心架构包含时序跨注意力预测模块、2D到3D视图变换器、3D编码器及用于多时序体素级预测的语义占位头。在两数据集上的广泛实验表明,ForecastOcc持续优于基线,生成具有丰富语义的前瞻预测,有效捕捉场景动态与语义,对自动驾驶至关重要。

原文摘要 · Abstract (English)

Autonomous driving requires forecasting both geometry and semantics over time to effectively reason about future environment states. Existing vision-based occupancy forecasting methods focus on motion-related categories such as static and dynamic objects, while semantic information remains largely absent. Recent semantic occupancy forecasting approaches address this gap but rely on past occupancy predictions obtained from separate networks. This makes current methods sensitive to error accumulation and prevents learning spatio-temporal features directly from images. In this work, we present ForecastOcc, the first framework for vision-based semantic occupancy forecasting that jointly predicts future occupancy states and semantic categories. Our framework yields semantic occupancy forecasts for multiple horizons directly from past camera images, without relying on externally estimated maps. We evaluate ForecastOcc in two complementary settings: multi-view forecasting on the Occ3D-nuScenes dataset and monocular forecasting on SemanticKITTI, where we establish the first benchmark for this task. We introduce the first baselines by adapting two 2D forecasting modules within our framework. Importantly, we propose a novel architecture that incorporates a temporal cross-attention forecasting module, a 2D-to-3D view transformer, a 3D encoder for occupancy prediction, and a semantic occupancy head for voxel-level forecasts across multiple horizons. Extensive experiments on both datasets show that ForecastOcc consistently outperforms baselines, yielding semantically rich, future-aware predictions that capture scene dynamics and semantics critical for autonomous driving.

自动驾驶语义占位多视角时序预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。