arXiv:2607.02712cs.CV2026-07中稿 · ECCV

让生成式视角合成实现任意角度精准控制,无需依赖输入图像姿态。

Global Pose Control for Generative View Synthesis in Normalized Object Coordinate Space

论文配图:Global Pose Control for Generative View Synthesis in Normalized Object Coordinate Space
图 1 · 摘自论文原文
  • 在归一化物体坐标系中直接操控绝对相机位姿,摆脱输入相对坐标限制。
  • 仅需1张或少量无姿态图像,即可生成多视角一致且高质量的新视图。
  • 结合文本描述定义物体标准坐标系,提升对未知类别物体的泛化能力。

新视角合成(NVS)可从单张或多张图像生成场景的未见视角,使用户能自由探索物体任意视角。尽管生成模型在定性效果上取得显著进展,现有方法仍难以实现全局、直观的视角控制,原因在于其依赖输入相对相机位姿,或仅能生成稀疏全局视图。这一局限严重制约了下游任务的扩展。为此,本文提出一种基于可定制归一化物体坐标空间(NOCS)的精确相机控制新方法,仅需单张或少数无姿态图像。该方法完全基于目标视图在NOCS中的绝对相机位姿,无需相对世界坐标系或输入图像的相机位姿。不同于以往将NVS视为独立生成任务的做法,本文将其建模为图像编辑问题,并利用先进的编辑模型以发挥其更强的泛化能力。通过上下文多模态条件策略注入相机信息作为专用标记。为缓解NOCS固有的模糊性,引入显式定义物体规范坐标系的文本描述,进一步提升对未见物体类别的泛化性能。此外,我们构建了一个高质量数据集,包含一致对齐的姿态和对应的NOCS文本定义。大量实验表明,本方法能从任意无姿态图像稳健生成具有准确且一致朝向的新视图,在多种物体类别上均达到当前最优的图像质量与保真度。

原文摘要 · Abstract (English)

Novel View Synthesis (NVS) enables the generation of unseen views of a scene from a single or multiple images, allowing users to freely explore an object from any viewpoint. Despite the recent impressive qualitative improvements of generative models for this task, existing methods struggle to provide global and intuitive control of target viewpoints because they either use input-relative camera poses or are limited to generating sparse global views. This lack of global pose control severely limits the number of downstream tasks potentially enabled by NVS. To address this limitation, we propose a novel approach for precise camera control in a customizable Normalized Object Coordinate Space (NOCS), requiring single or few unposed images. Our method operates solely on the absolute camera pose of the target view in NOCS, eliminating the need for a relative world frame or camera poses of the input images. Unlike previous methods that treat NVS as a standalone generation task, we formulate it as an image editing problem and build upon state-of-the-art editing models to leverage their superior generalization capability. Camera information is injected as dedicated camera tokens via an in-context multi-modal conditioning strategy. To alleviate the inherent ambiguity of NOCS, we incorporate text descriptions that explicitly define the object's canonical coordinate frame, which also enhances generalization to unseen object categories. Furthermore, we curate a high-quality dataset with consistently aligned orientations and corresponding NOCS text definitions. Extensive experiments demonstrate that our method robustly generates novel views with accurate and consistent orientations from arbitrary unposed images across diverse categories, achieving state-of-the-art image quality and fidelity.

视角合成姿态控制图像编辑归一化坐标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。