arXiv:2512.04830cs.CV2025-12

提出新框架,让自动驾驶场景生成更真实且视角间一致。

FreeGen: Feed-Forward Reconstruction-Generation Co-Training for Free-Viewpoint Driving Scene Synthesis

  • 重建与生成模型协同训练,提升视角一致性。
  • 在未见视角下生成效果优于现有方法,无需每场景优化。
  • 适合自动驾驶仿真与大规模训练的视觉生成需求。

封闭环路仿真与可扩展预训练需要合成自由视角的驾驶场景。然而,现有数据集和生成流程很少提供一致的非轨迹观测,限制了大规模评估与训练。尽管近期生成模型展现出强视觉真实感,但缺乏在不依赖场景微调的情况下同时实现插值一致性与外推真实感的能力。为此,我们提出 FreeGen,一种前馈式重建-生成协同训练框架,用于自由视角驾驶场景合成。重建模型提供稳定的几何表示以保证插值一致性,生成模型则进行感知几何的增强以提升未见视角的真实感。通过协同训练,生成先验被提炼至重建模型中,改善非轨迹渲染;同时,优化后的几何结构反过来为生成提供更强的结构引导。实验表明,FreeGen 在自由视角驾驶场景合成任务上达到当前最优性能。

原文摘要 · Abstract (English)

Closed-loop simulation and scalable pre-training for autonomous driving require synthesizing free-viewpoint driving scenes. However, existing datasets and generative pipelines rarely provide consistent off-trajectory observations, limiting large-scale evaluation and training. While recent generative models demonstrate strong visual realism, they struggle to jointly achieve interpolation consistency and extrapolation realism without per-scene optimization. To address this, we propose FreeGen, a feed-forward reconstruction-generation co-training framework for free-viewpoint driving scene synthesis. The reconstruction model provides stable geometric representations to ensure interpolation consistency, while the generation model performs geometry-aware enhancement to improve realism at unseen viewpoints. Through co-training, generative priors are distilled into the reconstruction model to improve off-trajectory rendering, and the refined geometry in turn offers stronger structural guidance for generation. Experiments demonstrate that FreeGen achieves state-of-the-art performance for free-viewpoint driving scene synthesis.

场景生成自动驾驶协同训练视角合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。