arXiv:2507.21983cs.LG2025-07被引 7

用强化学习提升广告文案生成效果,点击率提高6.7%。

Improving Generative Ad Text on Facebook using Reinforcement Learning

  • 通过历史广告表现数据做奖励信号,训练语言模型生成广告文案。
  • 大规模测试中点击率提升6.7%,显著改善广告投资回报。
  • 适合关注广告生成与真实效果优化的从业者参考。

生成式人工智能,特别是大语言模型(LLMs),正推动变革性经济影响。尽管LLMs在海量文本上预训练以掌握通用语言模式,但后续的后训练阶段对适配特定现实任务至关重要。强化学习(RL)是主流的后训练技术,但其经济影响仍缺乏系统探索与量化。本文通过首次在Facebook上部署强化学习训练的生成式广告模型展开研究。该模型“AdLlama”集成于Meta的文本生成功能,帮助广告商生成人类撰写广告文案的新变体。为训练此模型,我们提出基于性能反馈的强化学习(RLPF),利用历史广告表现数据作为奖励信号。在覆盖近3.5万广告商、64万广告变体的大规模10周A/B测试中,相较于基于精选广告微调的监督模型,AdLlama使点击率提升6.7%(p=0.0296),显著提升广告商投资回报。使用AdLlama的广告商还生成了更多广告变体,表明对输出结果满意度更高。据我们所知,这是迄今在生态有效环境下关于生成式AI应用的最大规模研究,为强化学习后训练的实际效益提供了关键量化依据。结果也表明,RLPF是一种有前景且可泛化的度量驱动后训练方法,能有效连接强大语言模型与实际产出。

原文摘要 · Abstract (English)

Generative artificial intelligence (AI), in particular large language models (LLMs), is poised to drive transformative economic change. LLMs are pre-trained on vast text data to learn general language patterns, but a subsequent post-training phase is critical to align them for specific real-world tasks. Reinforcement learning (RL) is the leading post-training technique, yet its economic impact remains largely underexplored and unquantified. We examine this question through the lens of the first deployment of an RL-trained LLM for generative advertising on Facebook. Integrated into Meta's Text Generation feature, our model, "AdLlama," powers an AI tool that helps advertisers create new variations of human-written ad text. To train this model, we introduce reinforcement learning with performance feedback (RLPF), a post-training method that uses historical ad performance data as a reward signal. In a large-scale 10-week A/B test on Facebook spanning nearly 35,000 advertisers and 640,000 ad variations, we find that AdLlama improves click-through rates by 6.7% (p=0.0296) compared to a supervised imitation model trained on curated ads. This represents a substantial improvement in advertiser return on investment on Facebook. We also find that advertisers who used AdLlama generated more ad variations, indicating higher satisfaction with the model's outputs. To our knowledge, this is the largest study to date on the use of generative AI in an ecologically valid setting, offering an important data point quantifying the tangible impact of RL post-training. Furthermore, the results show that RLPF is a promising and generalizable approach for metric-driven post-training that bridges the gap between highly capable language models and tangible outcomes.

广告生成强化学习大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。