arXiv:2412.05435cs.CV2024-12CVPR被引 116

UniScene统一生成驾驶场景的语义占据、视频和激光雷达数据。

UniScene: Unified Occupancy-centric Driving Scene Generation

论文配图:UniScene: Unified Occupancy-centric Driving Scene Generation
图 1 · 摘自论文原文
  • 以语义占据为中间表示,分两步生成多模态数据。
  • 在KITTI、nuScenes上生成质量优于现有SOTA方法。
  • 适合自动驾驶数据增强与多任务训练场景。

高保真、可控制且带标注的训练数据对自动驾驶至关重要。现有方法通常直接从粗粒度场景布局生成单一数据形式,不仅无法输出多样下游任务所需的丰富数据形态,也难以建模布局到数据的直接映射关系。本文提出UniScene,首个统一生成驾驶场景中语义占据、视频和激光雷达三种关键数据形式的框架。该框架采用渐进式生成流程,将复杂任务分解为两个层次步骤:(a) 从定制化场景布局生成语义占据作为富含语义与几何信息的元场景表示;(b) 基于占据表示,分别通过基于高斯的联合渲染与先验引导的稀疏建模两种新策略生成视频与激光雷达数据。这种以占据为中心的方法降低了生成负担,尤其在复杂场景中优势明显,并为后续生成阶段提供详细中间表示。大量实验表明,UniScene在语义占据、视频和激光雷达生成方面均超越现有SOTA方法,且显著提升下游驾驶任务性能。项目页面:https://arlo0o.github.io/uniscene/

原文摘要 · Abstract (English)

Generating high-fidelity, controllable, and annotated training data is critical for autonomous driving. Existing methods typically generate a single data form directly from a coarse scene layout, which not only fails to output rich data forms required for diverse downstream tasks but also struggles to model the direct layout-to-data distribution. In this paper, we introduce UniScene, the first unified framework for generating three key data forms - semantic occupancy, video, and LiDAR - in driving scenes. UniScene employs a progressive generation process that decomposes the complex task of scene generation into two hierarchical steps: (a) first generating semantic occupancy from a customized scene layout as a meta scene representation rich in both semantic and geometric information, and then (b) conditioned on occupancy, generating video and LiDAR data, respectively, with two novel transfer strategies of Gaussian-based Joint Rendering and Prior-guided Sparse Modeling. This occupancy-centric approach reduces the generation burden, especially for intricate scenes, while providing detailed intermediate representations for the subsequent generation stages. Extensive experiments demonstrate that UniScene outperforms previous SOTAs in the occupancy, video, and LiDAR generation, which also indeed benefits downstream driving tasks. Project page: https://arlo0o.github.io/uniscene/

场景生成自动驾驶多模态数据占据表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。