单阶段生成电商海报,同时精准控制主体、文字和风格。
InnoAds-Composer: Efficient Condition Composition for E-Commerce Poster Generation
- 用动态路由机制只在关键位置注入条件信息,降低计算开销。
- 中文文字渲染准确率提升,依赖图文特征融合模块。
- 首个包含三重条件的电商海报数据集,适合视觉生成研究者。
电商产品海报生成旨在通过主体、文字与设计风格合成一张有效传达商品信息的图像。近年来,具备细粒度可控性的扩散模型推动了海报合成进展,但多数方法依赖多阶段流程,对主体、文字与风格的协同控制仍不足。传统多阶段方案存在主体保真度差、文字不准、风格不一致等问题。为此,我们提出InnoAds-Composer,一个支持主体、字形(glyph)与风格三重条件的单阶段框架。为缓解朴素三条件拼接带来的二次复杂度,我们分析各层与时间步的重要性,仅将每类条件路由至最敏感位置,从而缩短有效激活序列。此外,为提升中文文字渲染精度,设计文本特征增强模块(TFEM),融合字形图像与裁剪图特征。为支持训练与评估,构建首个联合包含主体、文字与风格条件的高质量电商海报数据集与基准。大量实验表明,InnoAds-Composer显著优于现有方法,且推理延迟无明显增加。
原文摘要 · Abstract (English)
E-commerce product poster generation aims to automatically synthesize a single image that effectively conveys product information by presenting a subject, text, and a designed style. Recent diffusion models with fine-grained and efficient controllability have advanced product poster synthesis, yet they typically rely on multi-stage pipelines, and simultaneous control over subject, text, and style remains underexplored. Such naive multi-stage pipelines also show three issues: poor subject fidelity, inaccurate text, and inconsistent style. To address these issues, we propose InnoAds-Composer, a single-stage framework that enables efficient tri-conditional control tokens over subject, glyph, and style. To alleviate the quadratic overhead introduced by naive tri-conditional token concatenation, we perform importance analysis over layers and timesteps and route each condition only to the most responsive positions, thereby shortening the active token sequence. Besides, to improve the accuracy of Chinese text rendering, we design a Text Feature Enhancement Module (TFEM) that integrates features from both glyph images and glyph crops. To support training and evaluation, we also construct a high-quality e-commerce product poster dataset and benchmark, which is the first dataset that jointly contains subject, text, and style conditions. Extensive experiments demonstrate that InnoAds-Composer significantly outperforms existing product poster methods without obviously increasing inference latency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。