arXiv:2606.02595cs.LG2026-06被引 2

用人类审批机制让短租定价模型冷启动从150次缩短到30次。

Human-in-the-Loop Contextual Bandits for Short-Term Rental Dynamic Pricing: Structural Equivalence of Historical Warm-Up and Approval-Gated Live Learning

  • 人类审批约束下,历史数据可等效替代在线预热数据。
  • 真实数据验证:冷启动需求从约150次降至约30次。
  • 适用于医疗、信贷等高风险需人工审核的领域。

短租市场动态定价对在线学习算法提出独特挑战:定价决策财务风险高,运营者需要可解释性,且市场反馈稀疏(每晚仅一个预订结果)。本文提出人类在环门控强化学习框架(HITL-GB),上下文老虎机算法生成价格建议,由人工决定是否采纳、修改或拒绝。我们证明,在审批约束下,过往固定策略收集的历史定价数据,与同策略预热数据在结构上等价,可有效初始化老虎机后验分布,将纯在线学习中数周至数月的冷启动期大幅缩短。我们形式化了审批门控奖励信号,推导出基于正则化岭回归的预热方法,并在真实短租生产数据上验证(匿名城市市场,2间房,2022年4月至2026年4月,共1,461个夜间定价记录)。实验显示,当使用分层因子汤普森采样(HF-TS)家族代理时,有效冷启动从约150次压缩至约30次。进一步论证该等价性具有领域通用性:凡需法定或操作性人工审批的高风险场景——如临床用药剂量、信贷审批、内容审核、放射诊断——均满足相同条件并受益于该预热策略。因此,在监管行业,强制人工监督不仅是部署约束,更是统计资产。

原文摘要 · Abstract (English)

Dynamic pricing in short-term rental (STR) markets presents a distinctive challenge for online learning algorithms: pricing decisions carry significant financial risk, operators require explainability, and market feedback is sparse (one booking outcome per listed night). We introduce the Human-in-the-Loop Gated Bandit (HITL-GB) framework, in which a contextual bandit algorithm generates price recommendations but a human agent retains authority to accept, modify, or reject each recommendation before it is applied. We show that under this approval constraint, historical pricing data -- collected under a prior deterministic policy -- is structurally equivalent to on-policy warm-up data for initialising the bandit's posterior, bypassing the weeks-to-months cold-start period that renders pure online bandit learning impractical in sparse-feedback markets. We formalise the approval-gated reward signal, derive a regularised ridge-regression warm-up procedure from historical episodes, and validate the approach on real STR production data (anonymised urban market, 2 rooms, April 2022 -- April 2026, 1,461 nightly pricing episodes). Our warm-up procedure compresses effective cold-start from ~150 episodes to ~30 episodes when initialising agents from the Hierarchical Factored Thompson Sampling (HF-TS) family. We further argue that the structural equivalence result is domain-agnostic: any high-stakes domain where human approval is legally or operationally required -- including clinical drug dosing, credit origination, content moderation, and radiological diagnosis -- satisfies the same conditions and benefits from the same warm-up strategy. In regulated industries, mandatory human oversight is thus a statistical asset rather than a deployment constraint.

动态定价人类在环冷启动优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。