通过学习最优价格映射,实现更优的动态定价决策。
Harnessing Unimodality in Semiparametric Contextual Pricing via Oracle Price Map Learning
- 基于标量索引构建分层优化策略,自适应学习最优价格函数。
- 在高维线性与非参数场景下均实现近最优后悔上界。
- 适用于无分布假设的上下文环境,适合实际商业定价系统。
研究半参数标量索引估值模型下的上下文动态定价问题,其中潜在价值为 $v_t=μ_ ext{ast}(\mathsf c_t)+ξ_t$,包含未知效用函数 $μ_ ext{ast}$ 与未知噪声分布。关键决策对象是受标量索引 $u=μ_ ext{ast}(\mathsf c)$ 和噪声尾部决定的一维最优价格映射 $u\mapsto p^\ast(u)$。在噪声尾部函数 $β$-Hölder 光滑($β\geq 2$)且满足收入几何条件(保证唯一、稳定内点最大值)时,该最优映射本身具备 $(β-1)$-光滑性。本文提出 $\mathsf{ORBIT}$ 策略,采用粗到细的模块化设计:以标量预估索引为输入,定位每个活跃区间的基准价格,并在信任域内通过带宽凸优化学习局部多项式逼近。对于基础线性效用模型 $μ_ ext{ast}(\mathsf c)=\mathsf c^\topθ_ ext{ast}$,设计自适应椭圆探索机制,在不假设上下文分布的前提下在线生成所需标量预估。最终策略达到 $\widetilde{O}\big(T^{\frac{2β-1}{4β-3}}+\sqrt{dT}\big)$ 的后悔上界。固定维度 $d$ 时,建立匹配的下界,表明非参数最优映射学习项达到极小极大最优。相同标量预估接口亦可扩展至稀疏高维线性与非参数霍尔德效用情形。
原文摘要 · Abstract (English)
We study contextual dynamic pricing in a semiparametric scalar-index valuation model where the latent value is $v_t=μ_\ast(\mathsf c_t)+ξ_t$, with an unknown utility map $μ_\ast$ and an unknown additive noise distribution. The key decision object is the one-dimensional oracle price map $u\mapsto p^\ast(u)$ induced by the scalar index $u=μ_\ast(\mathsf c)$ and the noise tail. Under the $β$-Hölder smoothness of the tail function for $β\geq 2$ and a revenue-geometry condition that gives a unique, stable, interior maximizer, this oracle map is itself $(β-1)$-smooth. We exploit such structure through $\mathsf{ORBIT}$, a modular coarse-to-fine policy that takes a scalar pilot index as input, localizes a benchmark price in each active bin, and learns a local polynomial approximation of the oracle map inside a trust region via bandit convex optimization. For the baseline linear utility model $μ_\ast(\mathsf c)=\mathsf c^\topθ_\ast$, an adaptive elliptical exploration scheme constructs the required scalar pilot online without distributional assumptions on the contexts. The resulting policy achieves regret $\widetilde{O}\big(T^{\frac{2β-1}{4β-3}}+\sqrt{dT}\big)$. For fixed $d$, we establish a matching lower bound in the horizon dependence, unveiling that the nonparametric oracle-map learning term is minimax sharp. The same scalar-pilot interface also yields extensions to sparse high-dimensional linear utility and nonparametric Hölder utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。