用结构化控制修复跨数据集驾驶场景的视频渲染缺陷
SPVC: Structured and Panoptic Video Fixing for Cross-Dataset Driving Scene Rendering

- 结合相机位姿与地图信息,分步修复视频中的画面模糊与错位
- 在多个数据集上实现统一模型训练,减少对特定场景的依赖
- 适合自动驾驶仿真中需要稳定视觉输出的研究者使用
基于3D高斯点云的驾驶场景重建与渲染已成为自动驾驶模拟的重要组成部分。然而,在外推自车轨迹或编辑场景时,生成视角常出现模糊结构、时间闪烁及前景背景错位等问题。现有方法多针对特定场景,如单帧重绘或物体编辑修正。本文提出SPVC——一种面向跨数据集驾驶场景渲染的结构化与全景视频修复框架。该框架遵循四项设计原则:(1) 结构化修复,利用相机位姿、3D边界框和高清地图等显式空间条件引导修复,减少幻觉;(2) 全景修复,同时修正背景失真(如道路、建筑、车道)和前景车辆因编辑引入的外观不一致;(3) 视频修复,处理连续视频序列,利用时间线索进行纠错;(4) 跨数据集修复,仅需一个共享网络即可适配多个驾驶数据集,无需为每组数据单独训练。具体地,通过模拟低约束3DGS渲染与前景车辆插入伪影构建成对的劣化-清晰训练数据,并训练一个两阶段可控视频扩散模型,先修复视频级外观,再以结构化控制优化场景布局。
原文摘要 · Abstract (English)
Driving scene reconstruction and rendering, especially with 3D Gaussian Splatting, has become an important component of autonomous driving simulation. However, rendered views often degrade under extrapolated ego trajectories and scene edits, producing blurry structures, temporal flicker, and foreground-background misalignment. Existing refinement methods are commonly designed for a specific setting, such as image-level novel-view repair or object-editing correction. In this paper, we introduce SPVC, a structured and panoptic video fixing framework for cross-dataset driving scene rendering. The name summarizes four design principles. (1) Structured fixing denotes the use of explicit spatial conditions, including camera pose, 3D bounding boxes, and HD maps, to guide the repair process and reduce uncontrolled hallucination. (2) Panoptic fixing refers to correcting both background rendering artifacts, such as distorted roads, buildings, and lanes, and foreground vehicle artifacts introduced by scene editing, such as inconsistent object appearance. (3) Video fixing means that the model operates on driving sequences rather than isolated frames, allowing temporal cues to be used during artifact correction. (4) Cross-dataset fixing means that a single shared network is trained and applied across multiple driving datasets, reducing the need for dataset-specific or scene-specific fixers. Concretely, we construct paired degraded-clean training data by simulating under-constrained 3DGS rendering and foreground vehicle insertion artifacts, and train a two-stage controllable video diffusion model that first addresses video-level appearance and then refines scene layout with structured controls.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。