arXiv:2511.17221cs.CVcs.RO2025-11被引 3

用4D查询实现无需标注的3D语义占位学习,精度提升26%。

QueryOcc: Query-based Self-Supervision for 3D Semantic Occupancy

  • 通过跨帧采样的4D时空查询直接监督3D语义占位
  • 在Occ3D-nuScenes上达到26%的语义射线交并比提升
  • 支持视觉模型或激光雷达数据,适合自动驾驶场景

从图像中学习三维场景几何与语义是计算机视觉的核心挑战,也是自动驾驶的关键能力。由于大规模三维标注成本过高,近期工作探索直接从传感器数据中进行无监督学习。现有方法或依赖2D渲染一致性(3D结构隐式出现),或基于累积激光雷达点云的离散体素网格,限制了空间精度与可扩展性。我们提出QueryOcc,一种基于查询的自监督框架,通过独立采样相邻帧的4D时空查询,直接学习连续的3D语义占位。该框架可接受由视觉基础模型生成的伪点云或原始激光雷达数据作为监督信号。为实现长距离监督与恒定内存下的推理,引入一种收缩型场景表示,在保留近场细节的同时平滑压缩远场区域。QueryOcc在自监督Occ3D-nuScenes基准上,相比以往纯相机方法,语义射线交并比提升26%,且运行速度达11.6 FPS,证明直接4D查询监督能有效驱动强自监督占位学习。

原文摘要 · Abstract (English)

Learning 3D scene geometry and semantics from images is a core challenge in computer vision and a key capability for autonomous driving. Since large-scale 3D annotation is prohibitively expensive, recent work explores self-supervised learning directly from sensor data without manual labels. Existing approaches either rely on 2D rendering consistency, where 3D structure emerges only implicitly, or on discretized voxel grids from accumulated lidar point clouds, limiting spatial precision and scalability. We introduce QueryOcc, a query-based self-supervised framework that learns continuous 3D semantic occupancy directly through independent 4D spatio-temporal queries sampled across adjacent frames. The framework supports supervision from either pseudo-point clouds derived from vision foundation models or raw lidar data. To enable long-range supervision and reasoning under constant memory, we introduce a contractive scene representation that preserves near-field detail while smoothly compressing distant regions. QueryOcc surpasses previous camera-based methods by 26% in semantic RayIoU on the self-supervised Occ3D-nuScenes benchmark while running at 11.6 FPS, demonstrating that direct 4D query supervision enables strong self-supervised occupancy learning. https://research.zenseact.com/publications/queryocc/

3D占位自监督自动驾驶查询机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。