arXiv:2606.31258cs.CV2026-06

用3D物体先验补全视角扭曲中的缺失表面,解决极端视角合成失效问题。

WarpHammer: Densifying Scene Warps with 3D Object Priors for Extreme View Synthesis

论文配图:WarpHammer: Densifying Scene Warps with 3D Object Priors for Extreme View Synthesis
图 1 · 摘自论文原文
  • 利用3D生成先验补全物体遮挡区域的几何结构
  • 在视角偏移超过90度时仍能保持图像稳定与相机线索完整
  • 支持外部无位姿信息的物体参考图融合,无需微调模型

投影条件的新型视角合成(NVS)将输入视图的显式3D重建映射到目标相机,并以扭曲后的渲染结果作为生成器的条件。该方法在小视角变化下表现良好,但在大范围轨道运动下性能急剧下降:物体周围扭曲变得稀疏,隐藏表面主导新视角,出现镜面伪影,导致生成器丢失像素内容和扭曲中隐含的相机线索。我们提出WarpHammer,一种无需训练的框架,通过引入来自原生3D生成先验(如SAM3D)的显式物体3D重建来修复这一缺陷。该重建补充了缺失的前景表面并遮挡应不可见的背景点,恢复了外观和相机线索,且无需微调基础模型。同一显式物体表示还解锁了当前NVS流程不支持的能力:融合来自目标场景外的辅助视角,例如一张普通拍摄的汽车照片与厂商工作室拍摄的同款车型图像。我们联合处理参考图与辅助图像,借助预训练的多视角几何基础模型预测统一点云,并将其融合进3D物体重建中。这在不依赖辅助视角相机位姿的前提下,实现了比单图重建更忠实的几何结构。在五个基准测试中,WarpHammer在强基线崩溃的视角偏差条件下仍能生成稳定的新型视角,是首个可自然融合外部无位姿物体视图的场景级NVS方法。

原文摘要 · Abstract (English)

Projection-conditioned novel view synthesis (NVS) warps an explicit 3D reconstruction of the input view into the target camera and conditions a generator on the warped rendering. This works well for small viewpoint changes but degrades sharply under large orbital motion: the warp becomes sparse around the orbited object, where hidden surfaces dominate the new view and mirror-like artifacts emerge, causing the generator to lose both pixel content and the implicit camera cue carried by the warp. We introduce WarpHammer, a training-free framework that resolves this failure mode by augmenting the warped scene with an explicit 3D reconstruction of the object obtained from a native 3D generative prior (e.g., SAM3D). The reconstructed object adds missing foreground surfaces and occludes background points that should no longer be visible, restoring both appearance and camera cues without fine-tuning the base model. The same explicit object representation further unlocks a capability current NVS pipelines do not support: incorporating auxiliary views of the object from sources outside the target scene, for example, a casual snapshot of a car paired with a manufacturer studio shot of the same model. We process the reference and auxiliary images jointly with a pretrained multi-view geometry foundation model, which predicts a unified point cloud that we fuse into the 3D object reconstruction. This yields substantially more faithful geometry than single-image reconstruction, without requiring user-provided camera poses for the auxiliary views. On five benchmarks, WarpHammer produces stable novel views at viewpoint deviations where strong baselines collapse, and is the first scene-level NVS method that can naturally fuse auxiliary, pose-unknown object views from an external source.

视角合成3D先验物体重建多视图融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。