智能定价代理在信息不全时会误判市场,导致恶性竞争。
Market-Alignment Risk in Pricing Agents: Trace Diagnostics and Trace-Prior RL under Hidden Competitor State
- 通过分析历史市场轨迹,构建市场行为先验
- 修复后策略与对手在收益、房价和价格分布上高度一致
- 适合研究可解释性强化学习与市场对齐的开发者
在双酒店收益管理模拟器中,酒店A训练智能体对抗基于规则的酒店B。标准强化学习代理虽能达到接近参考的每间可用房收入(RevPAR),却未能学习到真实的市场收益管理:过度降价、自我打压或集中于低价区间。我们诊断此为部分可观测下的古德哈特式失效——酒店A无法观测对手剩余库存、预订曲线或定价规则,导致同一可见状态对应多个可能的对手价格。确定性值函数型强化学习和确定性复制会将这种不确定性压缩为捷径行为。本文提出一种基于轨迹的诊断协议,包含RevPAR、入住率、平均房价(ADR)、完整价格桶分布、L1/JS距离及种子级置信区间。验证修复方案为追踪先验强化学习(Trace-Prior RL):从滞后的市场轨迹中学习分布式市场先验,再以RevPAR为奖励、以与先验的KL散度为惩罚,训练随机定价策略。最终策略在种子级不确定范围内,与酒店B的RevPAR、入住率、ADR和价格分布完全对齐,同时仍优化酒店A自身收益。核心贡献并非新优化器或排行榜,而是一种可复现的智能体系统失败-修复范式,适用于标量奖励易被操纵、真实行为仅体现在轨迹中的场景。关键发现:更高精确动作准确率可能恶化整体轨迹对齐,当目标是分布对齐时。
原文摘要 · Abstract (English)
Outcome metrics can certify the wrong behavior. We study this failure in a two-hotel revenue-management simulator where Hotel A trains an agent against a fixed rule-based revenue-management competitor, Hotel B. A standard learning agent can obtain near-reference revenue per available room (RevPAR) while failing to learn market-like yield management: it sells too aggressively, undercuts, or collapses to modal price buckets. We diagnose this as a Goodhart-style failure under partial observability. Hotel A cannot observe the competitor's remaining inventory, booking curve, or pricing rule, so the same Hotel A-visible state maps to multiple plausible Hotel B prices. Deterministic value-based RL and deterministic copying collapse this unresolved uncertainty into shortcut behavior. We introduce a trace-level diagnostic protocol using RevPAR, occupancy, ADR, full price-bucket distributions, L1/JS distances, and seed-level confidence intervals. The verified repair is Trace-Prior RL: learn a distributional market prior from lagged market traces, then train a stochastic pricing policy with a RevPAR reward and a KL penalty to the learned prior. The final policy matches Hotel B's RevPAR, occupancy, ADR, and price distribution within seed-level uncertainty, while still optimizing Hotel A's own reward. We argue that the contribution is not a new optimizer and not a hotel-pricing leaderboard, but a reproducible failure-and-repair recipe for agentic systems where scalar rewards are easy to game and the intended behavior is only visible in traces. A key finding is that higher exact action accuracy can worsen aggregate trace alignment when the target is distributional.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。