arXiv:2509.04932cs.CV2025-09被引 1

用参考图提升单图新视角生成质量,减少扭曲

UniView: Enhancing Novel View Synthesis From A Single Image By Unifying Reference Features

论文配图:UniView: Enhancing Novel View Synthesis From A Single Image By Unifying Reference Features
图 1 · 摘自论文原文
  • 引入检索与增强系统,通过大模型选参考图
  • 设计多层隔离适配器动态生成参考特征
  • 解耦三重注意力保持输入细节,适合图像生成研究者

从单张图像合成新视角任务因未观察区域存在多种解释而高度病态。现有方法多依赖模糊先验和输入视角附近插值,常导致严重失真。为此,我们提出UniView模型,利用相似物体的参考图像提供强先验信息。具体而言,构建了检索与增强系统,并借助多模态大语言模型筛选满足需求的参考图像;引入可即插即用的适配器模块,通过多层级隔离层动态生成目标视角的参考特征;同时设计解耦三重注意力机制,有效对齐并融合多分支特征,保留原始输入细节。大量实验表明,UniView显著提升新视角生成性能,在挑战性数据集上优于当前最优方法。

原文摘要 · Abstract (English)

The task of synthesizing novel views from a single image is highly ill-posed due to multiple explanations for unobserved areas. Most current methods tend to generate unseen regions from ambiguity priors and interpolation near input views, which often lead to severe distortions. To address this limitation, we propose a novel model dubbed as UniView, which can leverage reference images from a similar object to provide strong prior information during view synthesis. More specifically, we construct a retrieval and augmentation system and employ a multimodal large language model (MLLM) to assist in selecting reference images that meet our requirements. Additionally, a plug-and-play adapter module with multi-level isolation layers is introduced to dynamically generate reference features for the target views. Moreover, in order to preserve the details of an original input image, we design a decoupled triple attention mechanism, which can effectively align and integrate multi-branch features into the synthesis process. Extensive experiments have demonstrated that our UniView significantly improves novel view synthesis performance and outperforms state-of-the-art methods on the challenging datasets.

新视角生成参考图像注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。