arXiv:2603.14965cs.CV2026-03

用3D几何引导生成更真实的新视角视频,解决画面扭曲问题。

GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis

  • 通过特征空间的高斯点云适配器实现几何约束
  • 相比现有方法,视角误差降低2倍,距离误差降7倍
  • 无需重训练即可适配多种3D模型,适合视频生成研究者

新视角合成需要强3D几何一致性与跨多视角的视觉连贯性。尽管近期相机控制的视频扩散模型表现良好,但仍存在几何失真和相机控制能力有限的问题。为此,我们提出GeoNVS,一种基于3D几何引导的新视角合成器,通过显式3D几何指导提升几何保真度和相机可控性。核心创新是高斯点云特征适配器(GS-Adapter),将输入视图的扩散特征提升至3D高斯表示,渲染出受几何约束的新视角特征,并自适应融合回扩散特征以修正几何不一致。与以往在输入层注入几何的方法不同,GS-Adapter在特征空间操作,避免了视图依赖的颜色噪声对结构一致性的影响。其即插即用设计支持零样本兼容多种前馈式几何模型,无需额外训练,且可适配其他视频扩散主干网络。在9个场景、18种设置下的实验表明,性能达到当前最优,相较SEVA和CameraCtrl分别提升11.3%和14.9%,翻译误差减少2倍,Chamfer Distance降低7倍。

原文摘要 · Abstract (English)

Novel view synthesis requires strong 3D geometric consistency and the ability to generate visually coherent images across diverse viewpoints. While recent camera-controlled video diffusion models show promising results, they often suffer from geometric distortions and limited camera controllability. To overcome these challenges, we introduce GeoNVS, a geometry-grounded novel-view synthesizer that enhances both geometric fidelity and camera controllability through explicit 3D geometric guidance. Our key innovation is the Gaussian Splat Feature Adapter (GS-Adapter), which lifts input-view diffusion features into 3D Gaussian representations, renders geometry-constrained novel-view features, and adaptively fuses them with diffusion features to correct geometrically inconsistent representations. Unlike prior methods that inject geometry at the input level, GS-Adapter operates in feature space, avoiding view-dependent color noise that degrades structural consistency. Its plug-and-play design enables zero-shot compatibility with diverse feed-forward geometry models without additional training, and can be adapted to other video diffusion backbones. Experiments across 9 scenes and 18 settings demonstrate state-of-the-art performance, achieving 11.3% and 14.9% improvements over SEVA and CameraCtrl, with up to 2x reduction in translation error and 7x in Chamfer Distance.

视频生成扩散模型3D重建新视角合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。