arXiv:2506.06316cs.IRcs.AI2025-06被引 2

用强化学习+大模型自动优化个性化营销的A/B测试

A Reinforcement-Learning-Enhanced LLM Framework for Automated A/B Testing in Personalized Marketing

  • 结合大模型生成内容,用强化学习动态选择最优广告版本
  • 在真实数据上提升点击率和转化率,优于传统方法
  • 适合需要持续优化广告策略的电商与推荐系统

针对个性化营销中如何高效自动化A/B测试以最大化用户响应的问题,本文提出一种RL-LLM-AB test框架。该框架基于预训练指令微调语言模型,首先通过提示条件生成器生成候选内容变体,再利用多模态感知模块融合用户画像与当前查询上下文,构建实时交互状态。随后,基于演员-评论家结构的策略优化模块实时选择内容版本,并根据点击率、转化率等反馈估计长期收益。此外,引入记忆增强型奖励估计算法,捕捉用户偏好漂移,提升跨用户与场景的策略泛化能力。在真实营销数据上的实验表明,该方法在点击率和转化率指标上显著优于经典A/B测试、上下文老虎机及基准强化学习方法。

原文摘要 · Abstract (English)

For personalized marketing, a new challenge of how to effectively algorithm the A/B testing to maximize user response is urgently to be overcome. In this paper, we present a new approach, the RL-LLM-AB test framework, for using reinforcement learning strategy optimization combined with LLM to automate and personalize A/B tests. The RL-LLM-AB test is built upon the pre-trained instruction-tuned language model. It first generates A/B versions of candidate content variants using a Prompt-Conditioned Generator, and then dynamically embeds and fuses the user portrait and the context of the current query with the multi-modal perception module to constitute the current interaction state. The content version is then selected in real-time through the policy optimization module with an Actor-Critic structure, and long-term revenue is estimated according to real-time feedback (such as click-through rate and conversion rate). Furthermore, a Memory-Augmented Reward Estimator is embedded into the framework to capture long-term user preference drift, which helps to generalize policy across multiple users and content contexts. Numerical results demonstrate the superiority of our proposed RL-LLM-ABTest over existing A/B testing methods, including classical A/B testing, Contextual Bandits, and benchmark reinforcement learning approaches on real-world marketing data.

A/B测试强化学习大模型应用个性化营销

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。