arXiv:2604.14059econ.GNcs.LG2026-04

对比动态规划与强化学习在有限期定价中的表现。

A Comparative Study of Dynamic Programming and Reinforcement Learning in Finite Horizon Dynamic Pricing

论文配图:A Comparative Study of Dynamic Programming and Reinforcement Learning in Finite Horizon Dynamic Pricing
图 1 · 摘自论文原文
  • 用数据估计需求,比较拟合动态规划与强化学习方法。
  • 在多产品、异质需求场景下,动态规划更稳定且约束满足更好。
  • 适合研究动态定价算法性能的学者或工业界优化人员。

本文系统比较了基于数据估计需求的拟合动态规划(Fitted DP)与强化学习(RL)方法在有限期动态定价问题中的表现。分析涵盖从单一类型基准到具有异质需求和跨时期收入约束的多类型设置,环境结构复杂性逐步提升。不同于将动态规划限制于低维场景的简化比较,本文将其应用于包含多个产品类型和约束的高维环境中。评估指标包括收益表现、稳定性、约束满足行为及计算可扩展性,揭示了基于期望的显式优化与基于轨迹的学习之间的权衡。

原文摘要 · Abstract (English)

This paper provides a systematic comparison between Fitted Dynamic Programming (DP), where demand is estimated from data, and Reinforcement Learning (RL) methods in finite-horizon dynamic pricing problems. We analyze their performance across environments of increasing structural complexity, ranging from a single typology benchmark to multi-typology settings with heterogeneous demand and inter-temporal revenue constraints. Unlike simplified comparisons that restrict DP to low-dimensional settings, we apply dynamic programming in richer, multi-dimensional environments with multiple product types and constraints. We evaluate revenue performance, stability, constraint satisfaction behavior, and computational scaling, highlighting the trade-offs between explicit expectation-based optimization and trajectory-based learning.

动态定价强化学习动态规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。