统一几何引导提升相机可控图像编辑的结构一致性
UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models
- 从表征、架构到损失函数三层注入统一几何信息
- 在多视角连续运动下显著减少几何漂移和结构退化
- 适合需要高精度三维一致性的图像编辑场景
相机可控图像编辑旨在合成给定场景在不同相机位姿下的新视图,同时严格保持跨视图几何一致性。然而,现有方法通常依赖零散的几何引导,例如仅在表示层面注入点云,且主要基于操作离散视图映射的图像扩散模型,这两方面限制共同导致连续相机运动下出现几何漂移和结构退化。我们观察到,尽管视频模型能提供连续视角先验,若几何引导仍零散,仍难以建立稳定几何理解。为此,我们提出UniGeo,通过在表征、架构和损失函数三个层面统一注入几何引导,系统性解决该问题。具体而言,在表征层面,引入帧解耦几何参考注入机制,提供鲁棒的跨视图几何上下文;在架构层面,设计几何锚点注意力以对齐多视图特征;在损失层面,提出轨迹终点几何监督策略,显式强化目标视图的结构保真度。在多个公开基准上的综合实验表明,无论在大范围还是小范围相机运动设置下,UniGeo在视觉质量与几何一致性上均显著优于现有方法。
原文摘要 · Abstract (English)
Camera-controllable image editing aims to synthesize novel views of a given scene under varying camera poses while strictly preserving cross-view geometric consistency. However, existing methods typically rely on fragmented geometric guidance, such as only injecting point clouds at the representation level despite models containing multiple levels, and are mainly based on image diffusion models that operate on discrete view mappings. These two limitations jointly lead to geometric drift and structural degradation under continuous camera motion. We observe that while leveraging video models provides continuous viewpoint priors for camera-controllable image editing, they still struggle to form stable geometric understanding if geometric guidance remains fragmented. To systematically address this, we inject unified geometric guidance across three levels that jointly determine the generative output: representation, architecture, and loss function. To this end, we propose UniGeo, a novel camera-controllable editing framework. Specifically, at the representation level, UniGeo incorporates a frame-decoupled geometric reference injection mechanism to provide robust cross-view geometry context. At the architecture level, it introduces geometric anchor attention to align multi-view features. At the loss function level, it proposes a trajectory-endpoint geometric supervision strategy to explicitly reinforce the structural fidelity of target views. Comprehensive experiments across multiple public benchmarks, encompassing both extensive and limited camera motion settings, demonstrate that UniGeo significantly outperforms existing methods in both visual quality and geometric consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。