让商品海报文字准确生成,尤其解决中文复杂字符难题
PosterMaker: Towards High-Quality Product Poster Generation with Accurate Text Rendering
- 用字级视觉特征控制文本渲染,提升准确性
- 文字渲染准确率超90%,产品细节还原度高
- 适合电商、广告等需要精准图文生成的场景
商品海报融合主体、场景与文字,是吸引客户的关键工具。使用现代图像生成方法制作海报具有价值,但主要挑战在于准确渲染文字,尤其是包含超过10,000个汉字的中文等复杂书写系统。本文识别出精确文字渲染的关键在于构建字符可区分的视觉特征作为控制信号。基于此,我们提出一种鲁棒的字级表示作为控制,并开发TextRenderNet,实现超过90%的文字渲染准确率。另一挑战是保持用户特定产品的真实感,为此我们引入基于图像修复的SceneGenNet,并提出主体保真度反馈学习以进一步提升还原度。结合TextRenderNet与SceneGenNet,我们构建了端到端的PosterMaker生成框架。为高效优化,采用两阶段训练策略,解耦文字渲染与背景生成的学习过程。实验表明,PosterMaker显著优于现有基线,验证了其有效性。
原文摘要 · Abstract (English)
Product posters, which integrate subject, scene, and text, are crucial promotional tools for attracting customers. Creating such posters using modern image generation methods is valuable, while the main challenge lies in accurately rendering text, especially for complex writing systems like Chinese, which contains over 10,000 individual characters. In this work, we identify the key to precise text rendering as constructing a character-discriminative visual feature as a control signal. Based on this insight, we propose a robust character-wise representation as control and we develop TextRenderNet, which achieves a high text rendering accuracy of over 90%. Another challenge in poster generation is maintaining the fidelity of user-specific products. We address this by introducing SceneGenNet, an inpainting-based model, and propose subject fidelity feedback learning to further enhance fidelity. Based on TextRenderNet and SceneGenNet, we present PosterMaker, an end-to-end generation framework. To optimize PosterMaker efficiently, we implement a two-stage training strategy that decouples text rendering and background generation learning. Experimental results show that PosterMaker outperforms existing baselines by a remarkable margin, which demonstrates its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。