用投影提示替代射线编码,提升视角合成的稳定性和一致性
From Rays to Projections: Better Inputs for Feed-Forward View Synthesis
- 用目标视图的投影提示替代射线参数作为输入
- 在基准测试中显著提升视角一致性和图像保真度
- 适合需要高几何一致性的视觉生成任务
前馈式视角合成模型通过单次前向传播生成新视角,但现有方法使用普吕克射线图作为相机编码,使预测结果依赖于任意的世界坐标系,对微小相机变换敏感,破坏几何一致性。本文提出投影条件化机制,将原始相机参数替换为稳定的二维目标视图投影提示,将原本脆弱的射线空间几何回归问题转化为良好的目标视图图像到图像翻译问题。同时设计了一种针对该提示的掩码自编码预训练策略,可利用大规模未标定数据进行预训练。在自建视角一致性基准上,本方法相比射线条件基线显著提升保真度与跨视角一致性,并在标准新视角合成基准上达到当前最优性能。
原文摘要 · Abstract (English)
Feed-forward view synthesis models predict a novel view in a single pass with minimal 3D inductive bias. Existing works encode cameras as Plücker ray maps, which tie predictions to the arbitrary world coordinate gauge and make them sensitive to small camera transformations, thereby undermining geometric consistency. In this paper, we ask what inputs best condition a model for robust and consistent view synthesis. We propose projective conditioning, which replaces raw camera parameters with a target-view projective cue that provides a stable 2D input. This reframes the task from a brittle geometric regression problem in ray space to a well-conditioned target-view image-to-image translation problem. Additionally, we introduce a masked autoencoding pretraining strategy tailored to this cue, enabling the use of large-scale uncalibrated data for pretraining. Our method shows improved fidelity and stronger cross-view consistency compared to ray-conditioned baselines on our view-consistency benchmark. It also achieves state-of-the-art quality on standard novel view synthesis benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。