用离线强化学习优化广告激励分配,平衡用户奖励与平台收益。
Learning to Allocate Incentives for Incentivized Advertising via Offline Model-Based Reinforcement Learning

- 构建用户反馈与广告收入的动态模型,实现序列化激励决策。
- 离线评估器可无成本预判策略效果,避免上线试错风险。
- 实测提升每用户净收益7.96%,适合广告平台精细化运营。
在激励式广告中,平台需提前承诺用户奖励以促使其观看并完成广告,但必须权衡预先承诺的激励与后续实际获得的广告收入:激励不足会错失变现机会,激励过量则降低净收益。由于当前激励会影响用户预期与未来参与度,激励分配成为具有延迟收益、成本敏感性及累积效应的序贯决策问题。现有研究未涉及此场景下的决策算法。自动出价依赖可用广告位,定向促销则独立于广告变现流程。本文将该问题建模为马尔可夫决策过程(MDP),提出一种离线模型增强型强化学习框架(Offline-MBRL),学习用户行为与广告收入的世界模型,并执行保守策略优化。引入独立反事实评分器,在保留日志上评估各策略,实现无需昂贵在线实验的预上线选择。大规模工业数据实验与在线A/B测试表明,该评分器提供稳定离线信号。从因果推断到离线强化学习再到离线MBRL的部署路径验证了框架有效性:相较于TD3+BC,MB-IQL提升每用户净收益7.96%;而退回普通IQL则下降6.56%(均p<0.0001)。
原文摘要 · Abstract (English)
Complete your ad view and grab a 5-cent bonus! In incentivized advertising, a platform promises users a bonus before observing downstream ad revenue, encouraging them to click and complete ads. It must balance the incentive promised in advance against the revenue realized afterward: insufficient incentives forfeit monetization opportunities, whereas excessive incentives reduce net profit. Because current incentives may also shape user expectations and future engagement, incentive allocation is a sequential decision problem with delayed revenue, cost sensitivity, and carryover effects. Existing work has not studied decision-making algorithms for this setting. Auto-bidding assumes available ad opportunities, while targeted promotion optimizes incentives outside the ad monetization pipeline. We formulate the problem as an MDP and develop an offline model-based RL framework for cost-controllable sequential incentive allocation. It learns a world model of user feedback and ad revenue, then performs conservative policy optimization. An independent counterfactual scorer evaluates each learned policy on held-out logs, enabling pre-launch selection without costly online exposure. Experiments on large-scale industrial data and online A/B tests show that the scorer provides a stable offline signal. The deployment path from causal inference to offline RL and then Offline-MBRL further validates the framework: MB-IQL improves per-user net profit by 7.96\% over TD3+BC, whereas reverting to plain IQL reduces it by 6.56\% (both \(p<0.0001\)).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。