设计公平且高效的积分奖励机制,兼顾客户差异与实验风险。
Learning Fair And Effective Points-Based Rewards Programs
- 用统一阈值设计公平方案,收入损失不超过1+ln2倍。
- 提出仅调整O(log T)次阈值的算法,实现最优近似后悔值。
- 阈值只降不升,提升公平性,代价仅为常数级后悔增加。
积分奖励计划是激励客户忠诚度的常见方式:客户通过重复购买积累积分,最终兑换免费奖励。然而,此类计划因实施中的不公平问题受到质疑。本文研究如何公平设计积分奖励机制,重点关注两个使公平与效果冲突的障碍:一是客户异质性要求差异化设置兑换门槛以提高收益;二是客户行为与积分积累关系未知,需通过实验探索,但可能不公平地贬低已有积分价值。我们首先证明,使用统一兑换阈值的个体公平方案,其收入损失最多为最优个性化策略的1+ln2倍。随后,针对需求不确定下的时间公平性问题,设计一种学习算法,仅在长度为T的周期内改变阈值O(log T)次,实现期望下最优(忽略多项式对数因子)的˜O(√T)后悔值。进一步修改该算法,使其仅降低阈值,从而提升公平性,仅带来常数级后悔增加。大量数值实验表明,在平均场景下个性化收益有限,同时验证了所提学习算法的强实用性。
原文摘要 · Abstract (English)
Points-based rewards programs are a prevalent way to incentivize customer loyalty; in these programs, customers who make repeated purchases from a seller accumulate points, working toward eventual redemption of a free reward. These programs have recently come under scrutiny due to accusations of unfair practices in their implementation. Motivated by these concerns, we study the problem of fairly designing points-based rewards programs, with a focus on two obstacles that put fairness at odds with their effectiveness. First, due to customer heterogeneity, the seller should set different redemption thresholds for different customers to generate high revenue. Second, the relationship between customer behavior and the number of accumulated points is typically unknown; this requires experimentation which may unfairly devalue customers' previously earned points. We first show that an individually fair rewards program that uses the same redemption threshold for all customers suffers a loss in revenue of at most a factor of $1+\ln 2$, compared to the optimal personalized strategy that differentiates between customers. We then tackle the problem of designing temporally fair learning algorithms in the presence of demand uncertainty. Toward this goal, we design a learning algorithm that limits the risk of point devaluation due to experimentation by only changing the redemption threshold $O(\log T)$ times, over a horizon of length $T$. This algorithm achieves the optimal (up to polylogarithmic factors) $\widetilde{O}(\sqrt{T})$ regret in expectation. We then modify this algorithm to only ever decrease redemption thresholds, leading to improved fairness at a cost of only a constant factor in regret. Extensive numerical experiments show the limited value of personalization in average-case settings, in addition to demonstrating the strong practical performance of our proposed learning algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。