arXiv:2601.21857cs.CV2026-01

用轨迹引导扩散模型生成文档背景,不破坏文字内容且保持多页风格一致。

Trajectory-Guided Diffusion for Foreground-Preserving Background Generation in Multi-Layer Documents

  • 在潜在空间设计扩散路径,通过初始噪声布局自然避开前景区域。
  • 引入缓存风格方向向量,实现跨页风格稳定一致,无需重复指定提示。
  • 无需训练、兼容现有模型,适合需要高保真文档生成的场景。

我们提出一种基于扩散模型的文档背景生成框架,通过潜在空间中的轨迹设计实现前景保留与多页风格一致性。该方法不依赖显式约束或掩码策略,而是将扩散过程重解为在结构化潜在空间中随机轨迹的演化。通过调整初始噪声及其几何对齐方式,背景生成可自然避开预定义的前景区域,确保文本可读性而不需额外机制。为解决跨页风格漂移问题,我们将风格控制与文本条件解耦,引入缓存风格方向作为潜在空间中的持久向量。一旦选定,这些方向将扩散路径约束于共享的风格子空间,确保多页及编辑迭代中外观一致。该框架无需训练,兼容现有扩散主干网络,能在复杂文档上生成视觉连贯、前景保留的结果。通过将扩散重新诠释为潜在流形上的轨迹设计,我们提供了一种原则性的统一生成建模方法。

原文摘要 · Abstract (English)

We present a diffusion-based framework for document-centric background generation that achieves foreground preservation and multi-page stylistic consistency through latent-space design rather than explicit constraints. Instead of suppressing diffusion updates or applying masking heuristics, our approach reinterprets diffusion as the evolution of stochastic trajectories through a structured latent space. By shaping the initial noise and its geometric alignment, background generation naturally avoids designated foreground regions, allowing readable content to remain intact without auxiliary mechanisms. To address the long-standing issue of stylistic drift across pages, we decouple style control from text conditioning and introduce cached style directions as persistent vectors in latent space. Once selected, these directions constrain diffusion trajectories to a shared stylistic subspace, ensuring consistent appearance across pages and editing iterations. This formulation eliminates the need for repeated prompt-based style specification and provides a more stable foundation for multi-page generation. Our framework admits a geometric and physical interpretation, where diffusion paths evolve on a latent manifold shaped by preferred directions, and foreground regions are rarely traversed as a consequence of trajectory initialization rather than explicit exclusion. The proposed method is training-free, compatible with existing diffusion backbones, and produces visually coherent, foreground-preserving results across complex documents. By reframing diffusion as trajectory design in latent space, we offer a principled approach to consistent and structured generative modeling.

文档生成扩散模型风格一致前景保留

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。