arXiv:2505.23434cs.CV2025-05

用分层语义几何先验提升城市场景的远超视图合成能力

UrbanCraft: Urban View Extrapolation via Hierarchical Sem-Geometric Priors

  • 构建粗粒度占用网格与细粒度3D框结合的分层先验
  • 在未见视角下合成效果显著优于传统扩散方法
  • 适合需要跨视角生成的城市重建与自动驾驶应用

现有基于神经渲染的城市场景重建方法主要针对训练相机轨迹附近的插值视图合成(IVS),但难以保证超出训练视角分布(如左右或向下看)的新视图质量,限制了泛化能力。此前方法尝试通过图像扩散改善,但因仅依赖文本控制,难以处理模糊文本或大角度未见视图。本文提出UrbanCraft,通过分层语义-几何先验解决外推视图合成(EVS)问题。具体地,利用部分可观测场景重建粗粒度语义与几何原型,以占用网格作为基础表示建立场景级先验;同时引入3D边界框信息提供实例级细节与空间关系。在此基础上,提出HSG-VSD方法,将预训练的UrbanCraft2D中的语义与几何约束融入得分蒸馏采样过程,强制生成分布与可观测场景一致。定性和定量实验均验证了该方法在EVS任务上的有效性。

原文摘要 · Abstract (English)

Existing neural rendering-based urban scene reconstruction methods mainly focus on the Interpolated View Synthesis (IVS) setting that synthesizes from views close to training camera trajectory. However, IVS can not guarantee the on-par performance of the novel view outside the training camera distribution (\textit{e.g.}, looking left, right, or downwards), which limits the generalizability of the urban reconstruction application. Previous methods have optimized it via image diffusion, but they fail to handle text-ambiguous or large unseen view angles due to coarse-grained control of text-only diffusion. In this paper, we design UrbanCraft, which surmounts the Extrapolated View Synthesis (EVS) problem using hierarchical sem-geometric representations serving as additional priors. Specifically, we leverage the partially observable scene to reconstruct coarse semantic and geometric primitives, establishing a coarse scene-level prior through an occupancy grid as the base representation. Additionally, we incorporate fine instance-level priors from 3D bounding boxes to enhance object-level details and spatial relationships. Building on this, we propose the \textbf{H}ierarchical \textbf{S}emantic-Geometric-\textbf{G}uided Variational Score Distillation (HSG-VSD), which integrates semantic and geometric constraints from pretrained UrbanCraft2D into the score distillation sampling process, forcing the distribution to be consistent with the observable scene. Qualitative and quantitative comparisons demonstrate the effectiveness of our methods on EVS problem.

场景重建视图外推生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。