arXiv:2606.22527cs.CV2026-06被引 1

让图像生成过程可编辑,从布局到细节分步可控。

Trajectory Forcing: Structure-First Generation with Controllable Semantic Trajectories

论文配图:Trajectory Forcing: Structure-First Generation with Controllable Semantic Trajectories
图 1 · 摘自论文原文
  • 按全局布局到细节的层级顺序生成图像,每步输出可查看可修改。
  • 在多个语义层级上实现结构一致性和局部控制,提升生成路径可读性。
  • 适合需要精细调控生成过程的研究者和设计师使用。

扩散模型和基于流的生成模型虽能生成高质量图像,但其可控性仍以最终结果为中心:用户设定条件后仅获得最终输出,中间生成过程不可见。现有方法开始利用生成顺序和过程分解提升样本质量,但仍将中间状态视为内部计算而非可交互对象。本文提出轨迹强迫(Trajectory Forcing, TF),一种以轨迹为核心的框架,使生成路径显式、语义化且可编辑。TF 将合成过程组织为一系列具有语义结构的阶段,依次从全局布局过渡到物体、部件和细节层级表示。每个阶段生成的潜在状态均可解码、检查、评估,并在进入下一阶段前进行局部编辑。为实现该路径,我们通过聚类预训练视觉表征(如 DINOv2)构建粗粒度到细粒度的教师层级结构,并在每个层级训练一个条件化的单步流匹配模型。此外,引入轨迹感知指标,衡量结构一致性与局部可控性,超越传统端点质量指标(如 FID)。实验表明,TF 在保持竞争性样本质量的同时,展现出连贯的中间状态,并支持跨语义层级的局部编辑。通过将关注点从最终图像转向生成路径本身,TF 开启了可控、轨迹感知图像生成的新路径。

原文摘要 · Abstract (English)

Diffusion and flow-based generative models produce strong images, yet their controllability remains largely endpoint-centric: users specify conditions and receive final outputs, while the intermediate generative dynamics remain hidden. Recent methods have begun to exploit generation order and process decomposition to improve sample quality, but still treat intermediate states as internal computation rather than objects for interaction. We propose Trajectory Forcing (TF), a trajectory-centric framework that makes the generation path explicit, semantic, and editable. TF organizes synthesis as a sequence of semantically structured stages, progressing from global layout to object-, part-, and detail-level representations. Each stage produces a decodable latent state that can be inspected, evaluated, and locally edited before the next stage begins. To instantiate this path, we derive coarse-to-fine teacher hierarchies by clustering pretrained visual representations such as DINOv2, and train a hierarchy-conditioned one-step flow-matching model at each level. We further introduce trajectory-aware metrics that measure structural consistency and local controllability beyond endpoint quality metrics such as FID. Experiments show that TF achieves competitive sample quality while exposing coherent intermediate states and supporting localized edits across semantic levels. By shifting the focus from final images to the generative path itself, TF opens a route toward controllable, trajectory-aware image synthesis.

图像生成可控生成路径可编辑扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。