用自然语言指令统一理解图像并生成风格一致的文本,无需手动设置字体等属性。
UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing
- 通过视觉语言模型理解指令与图像上下文,自动设计文本内容与布局。
- 在多个基准上达到当前最优效果,生成文本与图像风格高度一致。
- 适合需要自动化图文编辑的应用,如广告设计、智能排版。
随着图像生成技术的快速发展,利用自然语言指令进行视觉文本编辑受到越来越多关注。该任务的主要挑战在于充分理解指令和参考图像,从而生成与图像风格一致的视觉文本。以往方法通常需要复杂地指定文本内容和属性(如字体大小、颜色、布局),却忽视了与参考图像的风格一致性。为此,我们提出UM-Text,一个统一的多模态模型,用于基于自然语言指令的上下文理解与视觉文本编辑。具体地,我们引入视觉语言模型(VLM)处理指令与参考图像,使文本内容和布局能根据上下文信息精细设计。为生成准确且和谐的视觉文本图像,我们进一步提出UM-Encoder,自动融合多种条件信息的嵌入表示,其组合方式由VLM根据输入指令动态配置。训练时,我们设计区域一致性损失,在潜在空间与RGB空间对字形生成提供更有效的监督,并采用三阶段训练策略进一步提升性能。此外,我们构建了包含20万张图像的大型数据集UM-DATA-200K,覆盖多样化场景,用于模型训练。在多个公开基准上的定性与定量实验表明,本方法达到当前最优性能。
原文摘要 · Abstract (English)
With the rapid advancement of image generation, visual text editing using natural language instructions has received increasing attention. The main challenge of this task is to fully understand the instruction and reference image, and thus generate visual text that is style-consistent with the image. Previous methods often involve complex steps of specifying the text content and attributes, such as font size, color, and layout, without considering the stylistic consistency with the reference image. To address this, we propose UM-Text, a unified multimodal model for context understanding and visual text editing by natural language instructions. Specifically, we introduce a Visual Language Model (VLM) to process the instruction and reference image, so that the text content and layout can be elaborately designed according to the context information. To generate an accurate and harmonious visual text image, we further propose the UM-Encoder to combine the embeddings of various condition information, where the combination is automatically configured by VLM according to the input instruction. During training, we propose a regional consistency loss to offer more effective supervision for glyph generation on both latent and RGB space, and design a tailored three-stage training strategy to further enhance model performance. In addition, we contribute the UM-DATA-200K, a large-scale visual text image dataset on diverse scenes for model training. Extensive qualitative and quantitative results on multiple public benchmarks demonstrate that our method achieves state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。