arXiv:2510.13063cs.CVcs.AI2025-10被引 11

首个无需几何先验的自监督新视角合成模型,实现跨场景视角迁移。

True Self-Supervised Novel View Synthesis is Transferable

  • 通过成对姿态估计与输入输出增强,解耦相机姿态与场景内容。
  • 在多个场景间实现视角重渲染的可迁移性,优于现有无几何假设模型。
  • 适用于无3D先验的通用新视角生成,适合跨域视觉任务研究者。

本文指出,判断模型是否真正具备新视角合成(NVS)能力的关键标准是可迁移性:从一个视频序列中提取的任意姿态表示,能否用于重渲染另一个3D场景中的相同相机轨迹。我们分析了先前自监督NVS方法,发现其预测的姿态不具备可迁移性——同一组姿态在不同3D场景中产生不同的相机轨迹。为此,我们提出XFactor,首个无需几何先验的自监督新视角合成模型。XFactor结合成对姿态估计与简单的输入输出增强策略,联合实现相机姿态与场景内容的解耦,并促进几何推理。令人惊讶的是,XFactor在使用无约束的隐式姿态变量的情况下仍能实现可迁移性,且不依赖任何3D归纳偏置或多视图几何概念(如显式参数化为SE(3)群元素)。我们引入新指标量化可迁移性,通过大规模实验表明,XFactor显著优于现有无姿态假设的NVS Transformer,并通过探针实验验证隐式姿态与真实世界姿态高度相关。

原文摘要 · Abstract (English)

In this paper, we identify that the key criterion for determining whether a model is truly capable of novel view synthesis (NVS) is transferability: Whether any pose representation extracted from one video sequence can be used to re-render the same camera trajectory in another. We analyze prior work on self-supervised NVS and find that their predicted poses do not transfer: The same set of poses lead to different camera trajectories in different 3D scenes. Here, we present XFactor, the first geometry-free self-supervised model capable of true NVS. XFactor combines pair-wise pose estimation with a simple augmentation scheme of the inputs and outputs that jointly enables disentangling camera pose from scene content and facilitates geometric reasoning. Remarkably, we show that XFactor achieves transferability with unconstrained latent pose variables, without any 3D inductive biases or concepts from multi-view geometry -- such as an explicit parameterization of poses as elements of SE(3). We introduce a new metric to quantify transferability, and through large-scale experiments, we demonstrate that XFactor significantly outperforms prior pose-free NVS transformers, and show that latent poses are highly correlated with real-world poses through probing experiments.

新视角合成自监督学习可迁移性无几何先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。