arXiv:2606.03736stat.MLcs.LG2026-06

动态定价中,资源有限时自适应推理提升收益与稳定性。

Adaptive Inference for Resource-Constrained Dynamic Pricing

  • 根据资源约束提前检查价格可行性,动态调整定价策略。
  • 理论证明信息时钟线性增长,误差上界为O(log T),性能稳定可靠。
  • 适合资源受限场景的智能定价系统,尤其适用于高动态市场。

在有限销售周期内研究动态定价问题,当资源容量受限时,观测数据与收益均受制于预设价格。资源耗尽可能导致目标价格区间不可行,从而改变定价策略生成的实验分布。我们提出基于推断的重求解控制器,在当前协变量到达前检查目标带可行性并记录定价混合。目标保留型与平滑控制器以群体均值对几何结构为预部署输入;学习型重心重求解则估计预先声明组件核的稳态均消耗向量。在仿射绑定容量族上,精确输入的目标保留控制器分配质量 $t^{-γ}$,获得概率意义上的信息时钟阶为 $T^{1-γ}$,半径为 $O_pigrace{T^{-(1-γ)/2}igrace}$,并在暴露面奖励恒等条件下,有符号流体基准差距上界为 $O( ext{log }T + T^{1-γ})$。在外部仿射面条件下,预声明目标支撑及多项式误差支出指数大于1时,学习型重心重求解具有概率线性信息时钟和 $O( ext{log }T)$ 有符号差距上界;中心局部定价在松弛容量与全局最优下也达到相同阶数。无预留的精确输入、目标兼容平滑替代方案给出概率线性时钟,$O_p(T^{-1/2})$ 的确定性包络半径且覆盖率与报告概率趋近于1,有符号差距上界为 $O( ext{log}^2 T)$。边界结果揭示了物理支撑丧失时机理,并说明若仅依赖 $1/t$ 目标分支,则信息量仅为 $O_p(1)$。该策略在预设支撑与信息条件满足时报告区间,否则放弃输出。

原文摘要 · Abstract (English)

We study dynamic pricing over a finite selling horizon when limited resource capacity determines revenue and the observations available for inference at a prespecified price. Resource depletion can remove the target neighborhood from the feasible price set, changing the experiment generated by the pricing policy. We develop inference-aware re-solving controllers that check target-band feasibility before current covariates arrive and log the pricing mixture. Target-reserved and smooth controllers take population mean-pair geometry as a predeployment input; learned barycentric re-solving instead estimates stationary mean-consumption vectors of predeclared component kernels. On an affine binding-capacity family, an exact-input target-reserved controller assigning mass $t^{-γ}$ obtains an information clock of order $T^{1-γ}$ in probability, radius $O_p\{T^{-(1-γ)/2}\}$, and, under an exposed-face reward identity, a signed fluid-benchmark gap bounded above by $O(\log T+T^{1-γ})$. Under the exogenous affine-face condition, predeclared target support, and polynomial error spending with exponent greater than one, learned barycentric re-solving has a linear information clock in probability and an $O(\log T)$ signed-gap upper bound; centered local pricing has the same orders under slack capacity and global target optimality. An exact-input, target-compatible smooth alternative without reservation gives a linear clock in probability, an $O_p(T^{-1/2})$ deterministic-envelope radius with unconditional coverage and reporting probability tending to one, and an $O(\log^2 T)$ signed-gap upper bound. Boundary results show when physical support is lost and why a $1/t$ target branch yields only $O_p(1)$ information if it is the sole target-local source. The policy reports an interval when its prespecified support and information conditions hold and otherwise abstains.

动态定价资源约束自适应推理在线优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。