用生成模型动态分配激励,精准控制广告投入回报率。
Generative Optimization for Incentivized Advertising with Global Level Constraints

- 将激励分配建模为条件序列生成,结合用户历史与系统压力
- 在真实数据上使长期收入提升12.3%,违规率降低67%
- 单个模型适配多种回报率约束,无需重新训练
激励式广告通过发放金钱或虚拟奖励驱动用户参与,核心挑战在于在严格全局约束下优化连续激励额度。该问题因高频交互、延迟反馈及疲劳等非马尔可夫用户行为而复杂化,现有提升建模与约束强化学习方法效果受限。为此,我们提出GOAL,一种感知约束的生成框架,将激励分配建模为条件序列生成问题。GOAL基于用户历史与系统级全局压力直接生成激励值,并引入分层因果状态编码器,捕捉局部行为动态与长程依赖。为实现灵活约束控制,提出安全约束策略优化(SCPO),学习单一生成策略,在不重训练的前提下泛化于多种投资回报率(ROI)约束。在大规模真实数据及模拟疲劳环境中的实验表明,相比强基线,GOAL显著提升长期收入与用户留存,同时大幅降低ROI违规率。
原文摘要 · Abstract (English)
Incentivized advertising allocates monetary or virtual rewards to drive user engagement, where a key challenge is optimizing continuous incentive magnitudes under strict global constraints. This problem is complicated by high-frequency interactions, delayed feedback, and non-Markovian user dynamics such as fatigue, which limit the effectiveness of existing uplift modeling and constrained reinforcement learning approaches. To address these challenges, we propose GOAL, a constraint-aware generative framework that formulates incentive allocation as a conditional sequence generation problem. GOAL directly generates incentive magnitudes conditioned on user histories and system-level global pressure, and integrates a hierarchical causal state encoder to capture both local behavioral dynamics and long-range dependencies. To enable flexible constraint control, we introduce \textbf{S}afe \textbf{C}onstrained \textbf{P}olicy \textbf{O}ptimization (SCPO), which learns a single generative policy that generalizes across a spectrum of ROI constraints without retraining. Experiments on large-scale real-world data and a synthetic fatigue-aware environment show that GOAL improves long-term revenue and user retention while substantially reducing ROI violation rates compared to strong baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。