arXiv:2506.21276cs.CV2025-06被引 20

实现文字图像中逐词排版控制,提升生成精度与可编辑性

WordCon: Word-level Typography Control in Scene Text Rendering

  • 基于文本-图像对齐框架,利用视觉定位模型增强图文对应
  • 提出混合高效微调方法WordCon,支持多场景无缝集成
  • 采用潜空间掩码损失与联合注意力损失,实现单词间语义解耦

在生成图像中实现精确的逐词排版控制仍是一大挑战。为此,我们构建了一个新的词级可控场景文本数据集,并提出Text-Image Alignment(TIA)框架。该框架利用基础模型提供的文本与局部图像区域间的跨模态对应关系,增强文本到图像(T2I)模型的训练效果。此外,我们提出WordCon,一种混合参数高效微调(PEFT)方法,通过重参数化选择性关键参数,提升效率与可移植性,支持艺术化文本渲染、文本编辑及图像条件下的文本渲染等多种流程。为进一步增强可控性,我们在潜空间引入掩码损失,引导模型聚焦于图像中的文本区域;同时采用联合注意力损失,在特征层面提供监督,促进不同单词间的解耦。定性和定量结果均表明本方法优于现有技术。数据集与源代码将向学术界开放。

原文摘要 · Abstract (English)

Achieving precise word-level typography control within generated images remains a persistent challenge. To address it, we newly construct a word-level controlled scene text dataset and introduce the Text-Image Alignment (TIA) framework. This framework leverages cross-modal correspondence between text and local image regions provided by grounding models to enhance the Text-to-Image (T2I) model training. Furthermore, we propose WordCon, a hybrid parameter-efficient fine-tuning (PEFT) method. WordCon reparameterizes selective key parameters, improving both efficiency and portability. This allows seamless integration into diverse pipelines, including artistic text rendering, text editing, and image-conditioned text rendering. To further enhance controllability, the masked loss at the latent level is applied to guide the model to concentrate on learning the text region in the image, and the joint-attention loss provides feature-level supervision to promote disentanglement between different words. Both qualitative and quantitative results demonstrate the superiority of our method to the state of the art. The datasets and source code will be available for academic use.

文本生成排版控制图像生成参数高效微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。