用强化学习统一广告文案生成与效果优化,提升转化率。
RELATE: A Reinforcement Learning-Enhanced LLM Framework for Advertising Text Generation
- 通过策略学习将点击和转化等多维目标融入生成过程
- 在严格合规约束下实现点击转化率显著提升
- 适合需要高转化率且受政策限制的广告系统
在线广告中,广告文案对吸引用户参与和提升广告主价值至关重要。现有工业系统通常采用两阶段范式:先生成候选文案,再根据点击率(CTR)等线上指标进行对齐。这种分离导致优化目标错位、转化漏斗效率低,难以实现全局最优。为此,我们提出RELATE,一种基于强化学习的端到端框架,将生成与目标对齐统一于单一模型中。不同于传统解耦方式,RELATE通过策略学习直接在生成过程中整合性能与合规目标。为更准确反映广告主最终价值,我们引入以转化为核心的指标,并与合规约束共同作为多维奖励信号,使模型生成在政策约束下提升转化表现的高质量广告文案。大规模工业数据集上的实验表明,RELATE持续优于基线方法。在线部署于生产广告平台后,在严格政策约束下实现了点击转化率(CTCVR)的统计显著提升,验证了该框架的鲁棒性与实际有效性。
原文摘要 · Abstract (English)
In online advertising, advertising text plays a critical role in attracting user engagement and driving advertiser value. Existing industrial systems typically follow a two-stage paradigm, where candidate texts are first generated and subsequently aligned with online performance metrics such as click-through rate(CTR). This separation often leads to misaligned optimization objectives and low funnel efficiency, limiting global optimality. To address these limitations, we propose RELATE, a reinforcement learning-based end-to-end framework that unifies generation and objective alignment within a single model. Instead of decoupling text generation from downstream metric alignment, RELATE integrates performance and compliance objectives directly into the generation process via policy learning. To better capture ultimate advertiser value beyond click-level signals, We incorporate conversion-oriented metrics into the objective and jointly model them with compliance constraints as multi-dimensional rewards, enabling the model to generate high-quality ad texts that improve conversion performance under policy constraints. Extensive experiments on large-scale industrial datasets demonstrate that RELATE consistently outperforms baselines. Furthermore, online deployment on a production advertising platform yields statistically significant improvements in click-through conversion rate(CTCVR) under strict policy constraints, validating the robustness and real-world effectiveness of the proposed framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。