用生成模型+多实例奖励学习,优化电商广告创意组合。
Generative Modeling with Multi-Instance Reward Learning for E-commerce Creative Optimization
- 先生成多样创意组合,再用强化学习优化选择。
- 通过点击等反馈反推各元素贡献,提升奖励信号精度。
- 适合做电商广告优化的工程师和研究者参考。
在电商广告中,选择最具吸引力的创意组合(如标题、图片、亮点)对吸引用户注意力和促进转化至关重要。然而,现有方法通常单独评估创意元素,难以应对组合数量呈指数级增长的搜索空间。为此,我们提出一种名为 GenCO 的新框架,将生成建模与多实例奖励学习相结合。其统一的两阶段架构首先利用生成模型高效生成多样化的创意组合,并通过强化学习优化生成过程,实现有效探索与迭代。随后,针对用户反馈稀疏的问题,采用多实例学习模型将组合级奖励(如点击率)归因到各个创意元素,从而提供更精准的反馈信号,指导生成模型产出更具效果的组合。该方法已在主流电商平台上部署,显著提升了广告收入,验证了其实际价值。此外,我们还公开了一个大规模工业数据集,以推动该领域的进一步研究。
原文摘要 · Abstract (English)
In e-commerce advertising, selecting the most compelling combination of creative elements -- such as titles, images, and highlights -- is critical for capturing user attention and driving conversions. However, existing methods often evaluate creative components individually, failing to navigate the exponentially large search space of possible combinations. To address this challenge, we propose a novel framework named GenCO that integrates generative modeling with multi-instance reward learning. Our unified two-stage architecture first employs a generative model to efficiently produce a diverse set of creative combinations. This generative process is optimized with reinforcement learning, enabling the model to effectively explore and refine its selections. Next, to overcome the challenge of sparse user feedback, a multi-instance learning model attributes combination-level rewards, such as clicks, to the individual creative elements. This allows the reward model to provide a more accurate feedback signal, which in turn guides the generative model toward creating more effective combinations. Deployed on a leading e-commerce platform, our approach has significantly increased advertising revenue, demonstrating its practical value. Additionally, we are releasing a large-scale industrial dataset to facilitate further research in this important domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。