在价格数据不全时,用悲观与乐观策略优化离线动态定价。
A Tale of Two Cities: Pessimism and Opportunism in Offline Dynamic Pricing
- 基于需求单调性,对未观测价格的收益进行上下界估计。
- 提出悲观策略保底线收益,乐观策略追求潜在高回报,均在无覆盖场景下有效。
- 适用于风险偏好不同的企业,可直接映射到定价决策中。
我们研究历史数据无法覆盖全部价格空间的离线动态定价问题,此时最优价格可能完全未被观测,这在实践中很常见且在动态环境中尤为困难。现有离线强化学习方法通常依赖完整或部分覆盖,在此场景下表现不佳。本文提出一种非参数化部分识别框架,利用价格与需求的单调关系,对未观测价格的收益进行边界估计。在此框架下,构建两种动态决策规则:一种是最大化最坏情况收益的悲观策略,另一种是最小化最坏情况遗憾的乐观策略。二者专为序列无覆盖环境设计,非现有悲观离线强化学习或静态乐观方法的简单扩展。我们建立了两种策略的有限样本遗憾上界,当最优价格被覆盖时恢复标准速率,并量化了未覆盖带来的额外成本。还开发了高效算法,通过模拟和机票定价应用验证,方法在无覆盖设置下显著优于标准离线强化学习基线。管理层面,该框架将企业风险态度直接映射到定价策略:追求收入稳定与下行保护的企业应选悲观策略,愿承担适度风险以获取未探索价格潜在收益的企业则宜选乐观策略。
原文摘要 · Abstract (English)
We study offline dynamic pricing when historical data provide incomplete coverage of the price space such that some candidate prices, including the optimal one, may be entirely unobserved. This setting is common in practice and is especially difficult in dynamic environments. Existing offline reinforcement learning methods typically rely on full or partial coverage and can therefore perform poorly in such settings. We develop a nonparametric partial identification framework for offline dynamic pricing that exploits the monotonicity of demand in price to bound the value of unobserved prices. Within this framework, we formulate two dynamic decision rules: a pessimistic policy that maximizes worst-case revenue and an opportunistic policy that minimizes worst-case regret. These rules are tailored to a sequential no-coverage environment and are not direct extensions of existing pessimistic offline RL or static opportunistic approaches. We establish finite-sample regret bounds for both policies, recovering the standard rate when the optimal price is covered and quantifying the additional cost when it is not. We also develop efficient algorithms and show, through simulations and an airline ticket application, that our methods outperform standard offline RL baselines in no-coverage settings. Managerially, the framework provides a practical mapping from a firm's risk posture to its pricing policy: firms seeking revenue stability and downside protection should prefer the pessimistic policy, whereas firms willing to bear measured risk for potential gains from underexplored prices should prefer the opportunistic policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。