arXiv:2605.21263cs.LG2026-05

不依赖需求模型,仅用单点收益反馈动态调价并适应市场变化。

Nonparametric Learning and Earning with One-Point Feedback under Nonstationarity

论文配图:Nonparametric Learning and Earning with One-Point Feedback under Nonstationarity
图 1 · 摘自论文原文
  • 基于每期单一价格的收益数据,用梯度近似更新定价策略。
  • 引入周期重启机制,有效应对市场突变或渐变带来的非平稳性。
  • 适合缺乏先验模型、需实时响应市场的动态定价场景。

企业越来越依赖动态定价来应对不断变化的客户需求,但在许多实际应用中,只能观测到每个时期单一报价所产生的收益。与此同时,由于客户偏好、竞争格局或外部冲击的变化,市场条件可能逐渐或突然改变。这带来了两个相互关联的挑战:从有限反馈中学习收入-需求关系,并在变化的环境中调整定价决策。我们研究卖家如何在无特定需求参数形式假设的前提下,有效学习与获利。提出一种学习框架,利用每期一个观测值构造基于收益的梯度近似来更新价格。为应对环境变化,引入周期性重启机制,使过时信息被逐步淘汰。当非平稳程度未知时,进一步设计元学习层,自适应地在多种重启策略间权衡。提供性能保证,揭示累积收益损失相对于完全知情基准如何随时间跨度和市场波动幅度而变化。使用合成数据和真实世界数据的仿真实验验证了所提方法的有效性。

原文摘要 · Abstract (English)

Firms increasingly rely on dynamic pricing to respond to evolving customer demand, yet in many applications they observe only the revenue generated by a single posted price in each period. At the same time, market conditions may shift gradually or abruptly due to changes in customer preferences, competition, or external shocks. These features create two intertwined challenges: learning the revenue--demand relationship from limited feedback and adapting pricing decisions to a changing environment. We study how a seller can learn and earn effectively under these constraints, without assuming a specific parametric form for demand. We develop a learning framework that updates prices using revenue-based gradient approximations constructed from one observation per period. To address environmental changes, we incorporate a restarting mechanism that periodically refreshes the learning process so that outdated information is discounted. When the degree of nonstationarity is unknown, we further introduce a meta-learning layer to adaptively hedge across multiple restarting schedules. We provide performance guarantees for our approach, showing how cumulative revenue loss relative to a fully informed benchmark depends on both the time horizon and the magnitude of market variation. Simulation experiments using synthetic and real-world data illustrate the effectiveness of the proposed procedures.

动态定价非平稳性在线学习收益优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。