arXiv:2608.27893cs.CV2026-08

用可执行代码生成电商创意,让设计更可控、可编辑且质量更高。

CommerceVibe: Learning to Design E-Commerce Creatives as Executable Visual Code via Dual-Feedback Reinforcement Learning

论文配图:CommerceVibe: Learning to Design E-Commerce Creatives as Executable Visual Code via Dual-Feedback Reinforcement Learning
图 1 · 摘自论文原文
  • 将电商创意转为可运行的HTML/CSS代码,支持结构化生成。
  • 双反馈强化学习使创意在1300个测试案例中得分达94.0/100。
  • 适合需要高质量、可复用设计的电商平台和设计师使用。

高质量的电商创意对展示产品和传递营销信息至关重要。近期扩散模型虽能规模化生成视觉吸引人的图像,但其扁平化的位图输出常导致文字失真、产品细节不一致,需人工修正后才能部署。此外,缺乏显式结构使得生成内容难以编辑和复用,复杂的设计要求也难以转化为可验证的训练信号。为此,我们提出CommerceVibe,将创意表示为可执行的视觉代码,并将生成任务建模为条件性的HTML/CSS程序合成。给定产品图片、设计要求和商品信息,该模型可生成可渲染、可编辑、可复用的创意内容。我们引入双反馈强化学习:基于规则的反馈评估代码渲染后的文本可读性、产品可见性与布局合法性;视觉语言模型(VLM)提供视觉反馈,从六个感知与商业维度评估生成结果与输入规范的一致性。两者互补提升约束满足率与感知质量。我们在超过28,000个电商样本上对Qwen3.5-9B进行监督微调(SFT),随后进行双反馈强化学习。在1,300例基准测试中,优化后的CommerceVibe模型得分为94.0/100,优于仅使用SFT的版本(87.3),并超越多个外部强模型。五位电商设计专家的盲评进一步验证了改进效果。CommerceVibe支持可控、可编辑、可扩展的电商创意生产。

原文摘要 · Abstract (English)

High-quality e-commerce creatives are essential for presenting products and conveying marketing messages. Recent diffusion models enable scalable creative generation and produce visually compelling images, but their flattened raster outputs often contain distorted text and inconsistent product details, requiring refinement before deployment. Moreover, without explicit structure, the resulting creatives are difficult to edit and reuse, while complex design requirements remain challenging to encode as verifiable training signals. To address these challenges, we present CommerceVibe, which represents creatives as executable visual code and formulates generation as conditional HTML/CSS program synthesis. Given product images, design requirements, and product information, it produces renderable, editable, and reusable creatives. We further introduce dual-feedback reinforcement learning, in which rule-based feedback evaluates rendered programs for text readability, product visibility, and layout validity, while visual feedback from a vision-language model (VLM) assesses rendered creatives against input specifications across six perceptual and commercial dimensions. Together, these complementary feedback signals improve both constraint satisfaction and perception-dependent quality. We perform supervised fine-tuning (SFT) of Qwen3.5-9B on over 28,000 e-commerce examples, followed by dual-feedback reinforcement learning. On a 1,300-case benchmark, the optimized CommerceVibe model achieves a weighted score of 94.0/100, compared with 87.3 for the SFT-only variant, and outperforms strong external models. Blind evaluations by five e-commerce design experts further validate these improvements. CommerceVibe supports controllable, editable, and scalable e-commerce creative production.

电商创意可执行代码强化学习视觉生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。