arXiv:2510.20055cs.LGstat.ML2025-10NeurIPS

用强化学习建模延迟反馈,优化个性化广告出价策略

Learning Personalized Ad Impact via Contextual Reinforcement Learning under Delayed Rewards

  • 将广告出价建模为带延迟奖励的上下文马尔可夫决策过程
  • 理论证明算法近似最优,误差随客户数增长缓慢
  • 适合需要精准长期效果评估的广告平台与算法研究者

在线广告平台通过自动化拍卖机制连接广告主与潜在客户,需有效竞价策略以最大化收益。准确估计广告影响需考虑三个关键因素:延迟且长期的效果、累积性广告影响(如增强或疲劳)以及用户异质性。然而,此前研究常未能联合处理这些效应。为此,本文将广告出价建模为具有延迟泊松奖励的上下文马尔可夫决策过程(CMDP)。为实现高效估计,提出一种两阶段最大似然估计器,并结合数据分割策略,确保估计误差可控,依赖于第一阶段估计器的准确性。在此基础上,设计了一种强化学习算法,用于推导高效的个性化竞价策略。该方法达到近似最优的后悔界 $ ilde{O}(dH^2 oot{2}{T})$,其中 $d$ 为上下文维度,$H$ 为轮次数,$T$ 为客户数量。理论结果通过模拟实验验证。

原文摘要 · Abstract (English)

Online advertising platforms use automated auctions to connect advertisers with potential customers, requiring effective bidding strategies to maximize profits. Accurate ad impact estimation requires considering three key factors: delayed and long-term effects, cumulative ad impacts such as reinforcement or fatigue, and customer heterogeneity. However, these effects are often not jointly addressed in previous studies. To capture these factors, we model ad bidding as a Contextual Markov Decision Process (CMDP) with delayed Poisson rewards. For efficient estimation, we propose a two-stage maximum likelihood estimator combined with data-splitting strategies, ensuring controlled estimation error based on the first-stage estimator's (in)accuracy. Building on this, we design a reinforcement learning algorithm to derive efficient personalized bidding strategies. This approach achieves a near-optimal regret bound of $\tilde{O}{(dH^2\sqrt{T})}$, where $d$ is the contextual dimension, $H$ is the number of rounds, and $T$ is the number of customers. Our theoretical findings are validated by simulation experiments.

广告竞价强化学习延迟奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。