用大模型生成广告图,直接优化点击率,效果更好。
CTR-Driven Advertising Image Generation with Multimodal Large Language Models
- 用多模态大模型,以点击率为目标生成广告图
- 结合强化学习和产品特性,提升点击率与相关性
- 适合电商广告生成,可直接落地应用
网络数据中,广告图像对吸引用户注意力、提升广告效果至关重要。现有方法主要关注图像美学质量,可能导致在线表现不佳。为此,本文探索利用多模态大语言模型(MLLMs)生成广告图像,并以点击率(CTR)为核心目标进行优化。首先,构建针对性预训练任务,基于大规模电商多模态数据集,赋予MLLMs广告图像生成的初始能力。为进一步提升生成图像的点击率,提出一种新型奖励模型,通过强化学习(RL)微调预训练的MLLMs,联合利用多模态特征并准确反映用户点击偏好。同时,设计以产品为中心的偏好优化策略,确保微调后生成背景内容与产品特征一致,增强广告图像的整体相关性与有效性。大量实验表明,该方法在在线与离线指标上均达到领先水平。代码与预训练模型已开源:https://github.com/Chenguoz/CAIG。
原文摘要 · Abstract (English)
In web data, advertising images are crucial for capturing user attention and improving advertising effectiveness. Most existing methods generate background for products primarily focus on the aesthetic quality, which may fail to achieve satisfactory online performance. To address this limitation, we explore the use of Multimodal Large Language Models (MLLMs) for generating advertising images by optimizing for Click-Through Rate (CTR) as the primary objective. Firstly, we build targeted pre-training tasks, and leverage a large-scale e-commerce multimodal dataset to equip MLLMs with initial capabilities for advertising image generation tasks. To further improve the CTR of generated images, we propose a novel reward model to fine-tune pre-trained MLLMs through Reinforcement Learning (RL), which can jointly utilize multimodal features and accurately reflect user click preferences. Meanwhile, a product-centric preference optimization strategy is developed to ensure that the generated background content aligns with the product characteristics after fine-tuning, enhancing the overall relevance and effectiveness of the advertising images. Extensive experiments have demonstrated that our method achieves state-of-the-art performance in both online and offline metrics. Our code and pre-trained models are publicly available at: https://github.com/Chenguoz/CAIG.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。