arXiv:2511.22287cs.CV2025-11

无需训练,一键生成视角多变但内容一致的图像集。

Match-and-Fuse: Consistent Generation from Unstructured Image Sets

  • 将图像集建模为图,通过成对联合生成保持一致性。
  • 在无掩码、无监督下实现跨图像共享内容的统一生成。
  • 适合需要从杂乱照片中生成连贯视觉内容的创作者。

我们提出 Match-and-Fuse——一种零样本、无需训练的一致性图像集控制生成方法,适用于包含共同视觉元素但视角、拍摄时间与背景各异的非结构化图像集合。不同于针对单图或密集采样视频的方法,本框架采用集合到集合的生成范式:给定源图像集与用户提示,生成的新集合能保持跨图像共享内容的一致性。核心思想是将任务建模为图结构,每个图像为节点,每条边触发一对图像的联合生成。该设计将所有成对生成整合为统一框架,在局部一致性基础上保障全集全局连贯性。通过融合图像对间的内部特征,并基于密集输入对应关系实现,无需掩码或人工标注;同时利用文本到图像模型中涌现的先验,当多视角共享同一画布时,自动促进一致生成。Match-and-Fuse 在一致性与视觉质量上达到当前最佳表现,开启了从图像集合中进行内容创作的新可能。

原文摘要 · Abstract (English)

We present Match-and-Fuse - a zero-shot, training-free method for consistent controlled generation of unstructured image sets - collections that share a common visual element, yet differ in viewpoint, time of capture, and surrounding content. Unlike existing methods that operate on individual images or densely sampled videos, our framework performs set-to-set generation: given a source set and user prompts, it produces a new set that preserves cross-image consistency of shared content. Our key idea is to model the task as a graph, where each node corresponds to an image and each edge triggers a joint generation of image pairs. This formulation consolidates all pairwise generations into a unified framework, enforcing local consistency while ensuring global coherence across the entire set. This is achieved by fusing internal features across image pairs, guided by dense input correspondences, without requiring masks or manual supervision, and by leveraging an emergent prior in text-to-image models that encourages coherent generation when multiple views share a single canvas. Match-and-Fuse achieves state-of-the-art consistency and visual quality, and unlocks new capabilities for content creation from image collections.

图像生成一致性零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。