arXiv:2602.02283cs.LGstat.ML2026-02

用行为模型辅助强化学习,解决预订收入管理中反馈延迟问题。

Choice-Model-Assisted Q-learning for Delayed-Feedback Revenue Management

  • 用固定选择模型估算延迟反馈,修正强化学习目标。
  • 在参数变化时收益提升最高达12.4%,稳定场景下与基线无差异。
  • 模型假设错误时会降低收益,提示需谨慎使用行为模型。

针对收入管理中客户取消和修改信息延迟数日才可获取的问题,本文提出选择模型辅助强化学习:利用校准的离散选择模型作为部分世界模型,在决策时刻估算延迟收益成分。在固定模型部署下,证明表格Q-learning使用该模型估算的目标值可收敛至最优Q函数的$O(\varepsilon/(1-γ))$邻域,其中$\varepsilon$为模型误差,另含$O(t^{-1/2})$采样项。基于61,619条酒店预订数据的模拟器(1,088次独立运行)实验表明:(i) 在平稳环境下与成熟缓冲DQN基线无统计差异;(ii) 面对家族内参数变化时表现良好,10种情形中有5种经霍姆-邦弗伦尼校正后显著提升,最高达12.4%;(iii) 当选择模型结构错误时持续退化,收益下降1.4–2.6%。结果明确了部分行为模型在迁移下的增益边界与偏差风险。

原文摘要 · Abstract (English)

We study reinforcement learning for revenue management with delayed feedback, where a substantial fraction of value is determined by customer cancellations and modifications observed days after booking. We propose \emph{choice-model-assisted RL}: a calibrated discrete choice model is used as a fixed partial world model to impute the delayed component of the learning target at decision time. In the fixed-model deployment regime, we prove that tabular Q-learning with model-imputed targets converges to an $O(\varepsilon/(1-γ))$ neighborhood of the optimal Q-function, where $\varepsilon$ summarizes partial-model error, with an additional $O(t^{-1/2})$ sampling term. Experiments in a simulator calibrated from 61{,}619 hotel bookings (1{,}088 independent runs) show: (i) no statistically detectable difference from a maturity-buffer DQN baseline in stationary settings; (ii) positive effects under in-family parameter shifts, with significant gains in 5 of 10 shift scenarios after Holm--Bonferroni correction (up to 12.4\%); and (iii) consistent degradation under structural misspecification, where the choice model assumptions are violated (1.4--2.6\% lower revenue). These results characterize when partial behavioral models improve robustness under shift and when they introduce harmful bias.

强化学习收入管理延迟反馈选择模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。