分离语义与空间信息,提升单向新视角合成的图像质量。
Resolving Representation Ambiguity in Feedforward Novel View Synthesis Transformer via Semantic-Spatial Decoupling

- 将语义和空间信息分开展示,避免相互干扰。
- 在多种模型上实现稳定提升,无额外推理延迟。
- 适合追求高质量图像生成的研究者和开发者。
基于Transformer的模型推动了单向新视角合成(NVS)的发展。现有架构如GS-LRM和LVSM将语义信息(如RGB)与空间信息(如Plücker射线)混合在共享特征空间中。由于Plücker射线天然具有网格状空间结构,这种设计会导致空间偏差干扰外观表征,降低渲染保真度。为此,我们提出将单向NVS Transformer的表征解耦为独立的语义与空间令牌。该解耦设计在各自分支中显式保留语义与空间信息,同时通过共享注意力路由保持跨分支交互。在此基础上,我们引入可选的分类监督与双向调制:前者提供分支特异性训练信号,后者增强两分支间的交互。值得注意的是,基础解耦设计因架构特性几乎不引入额外推理延迟。所提方法在解码器仅用和编码器-解码器两类单向NVS模型中均取得一致改进。
原文摘要 · Abstract (English)
Transformer-based models have advanced feedforward novel view synthesis (NVS). Current architectures such as GS-LRM and LVSM mix semantic information (e.g., RGB) and spatial information (e.g., Plücker rays) into a shared feature space. Since Plücker rays naturally carry lattice-like spatial structure, these designs can make the spatial bias interfere with appearance representation and degrade rendering fidelity. To this end, we propose to decouple the representation of feedforward NVS transformers into separate semantic and spatial tokens. The decoupled design keeps semantic and spatial information explicit in their branches while preserving cross-branch interaction through shared attention routing. Built on this design, we introduce optional categorized supervision and bidirectional modulation: the former provides branch-specific training signals, while the latter improves interaction between the two branches. Notably, the base decoupled design introduces virtually zero additional inference latency due to its architectural design. The proposed designs achieve consistent improvements, demonstrating effectiveness across decoder-only and encoder-decoder feedforward NVS models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。