arXiv:2412.08243cs.CV2024-12TPAMI被引 10

通过解耦几何与时间上下文,提升摄像头3D语义占位预测的准确性。

Hierarchical Context Alignment with Disentangled Geometric and Temporal Modeling for Semantic Occupancy Prediction

论文配图:Hierarchical Context Alignment with Disentangled Geometric and Temporal Modeling for Semantic Occupancy Prediction
图 1 · 摘自论文原文
  • 分离几何与时间上下文,利用深度置信度和相机位姿对齐特征
  • 在SemanticKITTI和NuScenes上优于当前最优方法,语义一致性更强
  • 适合自动驾驶中复杂场景理解任务,尤其关注遮挡与模糊问题

基于摄像头的3D语义占位预测(SOP)对于从有限的2D图像观测中理解复杂三维场景至关重要。现有SOP方法通常通过聚合上下文特征来辅助占位表示学习,缓解遮挡或模糊问题。然而,这些方法常因不同帧间相同位置特征语义不一致导致对齐错误,进而引发不可靠的上下文融合与不稳定的表示学习。为此,我们提出一种新的分层上下文对齐范式(Hi-SOP),首先将几何与时间上下文解耦并分别对齐,再融合两者以增强占位预测可靠性。该方法包含:(I) 分离的几何与时间对齐,分别利用深度置信度和相机位姿作为先验进行特征匹配;(II) 基于语义一致性的全局对齐与几何-时间体的融合。实验表明,该方法在SemanticKITTI与NuScenes-Occupancy数据集上的语义场景补全任务中优于当前最优方法,在NuScenes上的激光雷达语义分割任务中也表现更优。

原文摘要 · Abstract (English)

Camera-based 3D Semantic Occupancy Prediction (SOP) is crucial for understanding complex 3D scenes from limited 2D image observations. Existing SOP methods typically aggregate contextual features to assist the occupancy representation learning, alleviating issues like occlusion or ambiguity. However, these solutions often face misalignment issues wherein the corresponding features at the same position across different frames may have different semantic meanings during the aggregation process, which leads to unreliable contextual fusion results and an unstable representation learning process. To address this problem, we introduce a new Hierarchical context alignment paradigm for a more accurate SOP (Hi-SOP). Hi-SOP first disentangles the geometric and temporal context for separate alignment, which two branches are then composed to enhance the reliability of SOP. This parsing of the visual input into a local-global alignment hierarchy includes: (I) disentangled geometric and temporal separate alignment, within each leverages depth confidence and camera pose as prior for relevant feature matching respectively; (II) global alignment and composition of the transformed geometric and temporal volumes based on semantics consistency. Our method outperforms SOTAs for semantic scene completion on the SemanticKITTI & NuScenes-Occupancy datasets and LiDAR semantic segmentation on the NuScenes dataset. The project website is available at https://arlo0o.github.io/hisop.github.io/.

语义占位视觉感知自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。