arXiv:2608.16289cs.CV2026-08

统一生成与编辑电商海报,支持文字块的增删改。

PosterText: Towards Unified Visual Text Generation and Editing for E-commerce Poster

论文配图:PosterText: Towards Unified Visual Text Generation and Editing for E-commerce Poster
图 1 · 摘自论文原文
  • 将文字块视为原子单元,统一处理生成与编辑任务。
  • 在多个数据集上优于现有方法,实现高质量海报生成与精准编辑。
  • 适合需要灵活设计电商海报的设计师与自动化系统使用。

自动化电商海报设计需兼顾高质量生成与灵活编辑能力。现有方法多聚焦于端到端生成或分阶段设计流程,难以实现对已有海报的精确修改。为此,本文提出文本块生成与编辑(Text Patch Generation and Editing)这一统一任务范式,将文本块作为基本单位,涵盖海报生成、添加、删除与修改四种操作,并支持参考引导的风格控制。基于此,我们构建了PosterText模型,采用四阶段课程训练:文本渲染预训练、指令跟随训练、偏好对齐强化学习,以及空间引导自蒸馏以优化执行精度。同时,我们创建了一个大规模带块级标注的数据集和全面评估基准。大量实验表明,PosterText在生成与编辑性能上均达到领先水平,验证了该框架的有效性。

原文摘要 · Abstract (English)

Automated e-commerce poster design requires both high-quality poster generation and flexible editing of existing designs. However, most existing methods either target end-to-end poster generation or follow multi-stage design pipelines, with limited capability for flexible and precise editing of existing posters. To enable unified generation and editing of e-commerce posters, we introduce Text Patch Generation and Editing, a unified task formulation that treats text patches as atomic units and covers four operations: poster generation, patch addition, patch deletion, and patch modification, with optional reference-guided style control. Based on this, we propose PosterText, a unified model trained with a four-stage curriculum, including text rendering pretraining, instruction-following training, reinforcement learning for preference alignment, and spatial guidance self-distillation for execution refinement. We further construct a large-scale dataset with patch-level annotations and a comprehensive benchmark for evaluation. Extensive experiments demonstrate that PosterText achieves competitive performance against existing generation and editing approaches, validating the effectiveness of the proposed framework.

海报生成文本编辑统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。