arXiv:2605.19949cs.CV2026-05被引 1

用稀疏航拍图重建城市3D场景,避免鬼影和变形。

Feed-Forward Gaussian Splatting from Sparse Aerial Views

论文配图:Feed-Forward Gaussian Splatting from Sparse Aerial Views
图 1 · 摘自论文原文
  • 先建可靠结构的几何隐变量,再用补全令牌修复弱约束区域。
  • 单次前向传播完成重建,生成视角一致且细节丰富的城市场景。
  • 适合做城市级3D重建的工程师和研究人员参考。

从稀疏航拍图像重建大规模城市场景是一项关键但具有挑战性的任务。由于俯视和浅斜视角导致的拍摄姿态偏差,稀疏航拍数据存在显著的证据不平衡:屋顶和开阔区域被重复观测,而立面、远处建筑和遮挡结构缺乏多视角支持。现有前馈式3D高斯点云方法直接从稀疏输入回归确定性表示,常导致鬼影、融化的立面和拉伸的纹理。近期基于伪视图和视频的生成重建方法虽引入额外监督或生成先验,但往往难以区分真实观测与先验驱动内容,产生看似合理却不一致的结构。本文提出AnyCity,一种基于观测的生成重建框架,用于稀疏航拍城市场景。该框架首先预测一个受观测支持的几何隐变量以锚定可靠结构,随后利用支架条件化的航拍补全标记,生成门控残差更新以处理弱约束内容,最后进行高斯解码。训练阶段采用密集到稀疏的知识蒸馏,传递结构线索;同时通过门控标记条件的航拍适配视频扩散先验,提供精细的城市外观线索。观测保持目标确保重构结果与输入支持的几何一致。推理时,AnyCity仅需一次前向传播即可从稀疏航拍图重建最终3D高斯场景,实现秒级推理下的连贯城市新视角合成。在合成数据、航拍域、无人机纹理及真实场景上的实验表明,该方法在多个基准上持续优于前馈基线。

原文摘要 · Abstract (English)

Reconstructing large-scale urban scenes from sparse aerial views is a crucial yet challenging task. Due to biased top-down and shallow-oblique camera poses, sparse aerial captures exhibit strong evidence imbalance: roofs and open regions are repeatedly observed, while facades, distant buildings, and occluded structures receive little multi-view support. Existing feed-forward 3D Gaussian Splatting methods directly regress a deterministic representation from sparse inputs, but this often leads to ghosting, melted facades, and stretched textures. Recent pseudo-view and video-based generative reconstruction methods use additional supervision or generative priors. However, they often lack a clear separation between observed geometry and prior-driven content, which can lead to plausible but inconsistent structures. We propose AnyCity, an observation-grounded generative reconstruction framework for sparse aerial urban scenes. AnyCity first predicts an observation-supported geometry latent to anchor reliable structures, and then uses scaffold-conditioned aerial completion tokens to predict a gated residual update for weakly constrained content before Gaussian decoding. During training, dense-to-sparse distillation transfers structural cues from dense-view reconstruction, while an aerial-adapted video diffusion prior provides fine-grained urban appearance cues through gated token conditioning. Observation-preserving objectives keep the refined representation consistent with input-supported geometry. At inference time, AnyCity reconstructs the final 3D Gaussian scene from sparse aerial views in a single feed-forward pass, achieving coherent urban novel-view synthesis with second-level inference. Experiments on synthetic, aerial-domain, UAV-textured, and real-world scenes show consistent improvements over feed-forward baselines.

3D重建高斯点云城市建模生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。