arXiv:2512.17151cs.CV2025-12中稿 · ECCV被引 1

让文档背景自动生成且可编辑,文字始终清晰可读。

Text-Conditioned Background Generation for Editable Multi-Layer Documents

  • 用扩散模型隐空间掩码保护文字区域,防止被干扰。
  • 自动添加半透明底框,满足网页可读性标准(WCAG 2.2)。
  • 多页保持主题连贯,支持分层编辑与风格定制。

我们提出一种面向文档的背景生成框架,支持多页编辑与主题一致性。为确保文字可读性,采用受物理与优化中光滑屏障函数启发的隐空间掩码机制,软性抑制扩散过程中的更新。此外引入自动化可读性优化(ARO),在文字区域自动添加半透明圆角衬底,根据底层背景动态调整最小不透明度以满足WCAG 2.2感知对比度标准,实现无需人工干预的可读性保障与视觉和谐。通过摘要-指令递归流程,将每页压缩为紧凑表征,延续人类记忆式上下文保持机制,确保视觉元素在整篇文档中连贯演进。方法将文档视为结构化组合,文本、图表与背景作为独立图层保留或再生,支持针对性背景编辑而不影响可读性。用户提示可灵活调整色彩与纹理,在自动化一致性与个性化定制间取得平衡。本训练无关框架生成视觉连贯、文字保真、主题一致的文档,弥合生成建模与自然设计工作流之间的鸿沟。

原文摘要 · Abstract (English)

We present a framework for document-centric background generation with multi-page editing and thematic continuity. To ensure text regions remain readable, we employ a latent masking formulation that softly attenuates updates in the diffusion space, inspired by smooth barrier functions in physics and numerical optimization. In addition, we introduce Automated Readability Optimization (ARO), which automatically places semi-transparent, rounded backing shapes behind text regions. ARO determines the minimal opacity needed to satisfy perceptual contrast standards (WCAG 2.2) relative to the underlying background, ensuring readability while maintaining aesthetic harmony without human intervention. Multi-page consistency is maintained through a summarization-and-instruction process, where each page is distilled into a compact representation that recursively guides subsequent generations. This design reflects how humans build continuity by retaining prior context, ensuring that visual motifs evolve coherently across an entire document. Our method further treats a document as a structured composition in which text, figures, and backgrounds are preserved or regenerated as separate layers, allowing targeted background editing without compromising readability. Finally, user-provided prompts allow stylistic adjustments in color and texture, balancing automated consistency with flexible customization. Our training-free framework produces visually coherent, text-preserving, and thematically aligned documents, bridging generative modeling with natural design workflows.

文档生成可读性优化多页一致扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。