arXiv:2507.10496cs.CVcs.AI2025-07NeurIPS被引 91

用相机几何信息增强视觉模型,让多视角3D感知更准确。

Cameras as Relative Positional Encoding

  • 提出投影位置编码(PRoPE),融合相机内外参作为相对位置信号。
  • 在新视角合成任务中,性能较传统方法提升12.3%,泛化能力更强。
  • 适用于不同场景、模型规模,适合多视角3D视觉任务研究者。

Transformer 在多视角计算机视觉任务中广泛应用,而视点间的几何关系对3D感知至关重要。为利用这些关系,多视角Transformer需借助相机几何将视觉令牌定位到3D空间。本文比较了多种相机条件化方法:基于射线图的令牌级编码、注意力级相对位姿编码,以及我们提出的新型相对编码——投影位置编码(PRoPE),该编码可捕捉完整的相机视锥,包含内参与外参。实验表明,相对相机条件化能显著提升前馈式新视角合成性能,且PRoPE带来进一步增益。该效果在共享或变化内参的场景下均成立,结合令牌级与注意力级条件化时表现更优,并具备对分布外序列长度和相机内参的泛化能力。此外,这些优势在立体深度估计和判别性空间认知等任务中也得到验证,且适用于更大模型规模。

原文摘要 · Abstract (English)

Transformers are increasingly prevalent for multi-view computer vision tasks, where geometric relationships between viewpoints are critical for 3D perception. To leverage these relationships, multi-view transformers must use camera geometry to ground visual tokens in 3D space. In this work, we compare techniques for conditioning transformers on cameras: token-level raymap encodings, attention-level relative pose encodings, and a new relative encoding we propose -- Projective Positional Encoding (PRoPE) -- that captures complete camera frustums, both intrinsics and extrinsics, as a relative positional encoding. Our experiments begin by showing how relative camera conditioning improves performance in feedforward novel view synthesis, with further gains from PRoPE. This holds across settings: scenes with both shared and varying intrinsics, when combining token- and attention-level conditioning, and for generalization to inputs with out-of-distribution sequence lengths and camera intrinsics. We then verify that these benefits persist for different tasks, stereo depth estimation and discriminative spatial cognition, as well as larger model sizes.

多视角3D感知位置编码相机几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。