arXiv:2505.19291cs.CVcs.AI2025-05

用强化学习加速文本布局生成,让图文合成更快更省资源。

TextDiffuser-RL: Efficient and Robust Text Layout Optimization for High-Fidelity Text-to-Image Synthesis

  • 两阶段流程:强化学习快速生成文本框,再用扩散模型绘图
  • 比TextDiffuser-2快42.29倍,仅需2MB CPU内存推理
  • 适合需要低延迟、跨平台部署的图形设计与内容生成场景

嵌入文本的图像生成在平面设计、广告和数字内容创作中至关重要。基于扩散模型的TextDiffuser-2等方法在生成带文本图像方面表现良好,能有效生成引导视觉文本渲染的边界框布局,实现高保真与一致性。然而现有方法依赖高资源消耗过程,难以在CPU或GPU上高效运行。为此,本文提出一种新型两阶段框架,将强化学习(RL)用于快速优化文本布局生成,并与扩散模型结合。该方法显著加速边界框预测并减少重叠,可在CPU与GPU上高效运行。大量实验表明,本框架在文本定位与图像合成质量上与TextDiffuser-2相当,但推理速度提升42.29倍,且仅需2MB CPU内存,而TextDiffuser-2的M1模型无法在纯CPU系统运行。

原文摘要 · Abstract (English)

Text-embedded image generation plays a critical role in industries such as graphic design, advertising, and digital content creation. Text-to-Image generation methods leveraging diffusion models, such as TextDiffuser-2, have demonstrated promising results in producing images with embedded text. TextDiffuser-2 effectively generates bounding box layouts that guide the rendering of visual text, achieving high fidelity and coherence. However, existing approaches often rely on resource-intensive processes and are limited in their ability to run efficiently on both CPU and GPU platforms. To address these challenges, we propose a novel two-stage pipeline that integrates reinforcement learning (RL) for rapid and optimized text layout generation with a diffusion-based image synthesis model. Our RL-based approach significantly accelerates the bounding box prediction step while reducing overlaps, allowing the system to run efficiently on both CPUs and GPUs. Extensive evaluations demonstrate that our framework achieves comparable performance to TextDiffuser-2 in terms of text placement and image synthesis, while offering markedly faster runtime and increased flexibility. Our method produces high-quality images comparable to TextDiffuser-2, while being 42.29 times faster and requiring only 2 MB of CPU RAM for inference, unlike TextDiffuser-2's M1 model, which is not executable on CPU-only systems.

图文生成强化学习扩散模型高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。