arXiv:2510.14981cs.CVcs.AI2025-10被引 6

用隐式3D约束实现无需训练的多视角图像一致编辑

Coupled Diffusion Sampling for Training-Free Multi-View Image Editing

  • 通过耦合扩散采样同步生成多视角图像
  • 在三种任务上实现跨视角一致性,效果优于显式3D优化
  • 适合作为通用多视角编辑框架,兼容多种模型架构

我们提出一种推理阶段的扩散采样方法,利用预训练的2D图像编辑模型实现多视角图像的一致编辑。这些模型虽能独立生成高质量单视角编辑结果,但无法保证跨视角一致性。现有方法通常依赖显式3D表示进行优化,但存在计算耗时长、稀疏视角下不稳定等问题。本文提出一种隐式3D正则化方法,通过约束生成的2D图像序列符合预训练的多视角图像分布来实现一致性。该方法采用耦合扩散采样,同时从多视角图像分布和2D编辑图像分布中采样两条轨迹,并通过耦合项强制生成图像间的多视角一致性。我们在三个不同的多视角图像编辑任务上验证了该框架的有效性与通用性,证明其可适配多种模型架构,展现出作为通用多视角一致编辑解决方案的潜力。

原文摘要 · Abstract (English)

We present an inference-time diffusion sampling method to perform multi-view consistent image editing using pre-trained 2D image editing models. These models can independently produce high-quality edits for each image in a set of multi-view images of a 3D scene or object, but they do not maintain consistency across views. Existing approaches typically address this by optimizing over explicit 3D representations, but they suffer from a lengthy optimization process and instability under sparse view settings. We propose an implicit 3D regularization approach by constraining the generated 2D image sequences to adhere to a pre-trained multi-view image distribution. This is achieved through coupled diffusion sampling, a simple diffusion sampling technique that concurrently samples two trajectories from both a multi-view image distribution and a 2D edited image distribution, using a coupling term to enforce the multi-view consistency among the generated images. We validate the effectiveness and generality of this framework on three distinct multi-view image editing tasks, demonstrating its applicability across various model architectures and highlighting its potential as a general solution for multi-view consistent editing.

多视角编辑扩散模型图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。