arXiv:2504.03743cs.LGcs.AI2025-04中稿 · RLDM 2025被引 2

用沃瑟斯坦距离建模有限理性决策,更好捕捉动作间的远近关系。

Modelling bounded rational decision-making through Wasserstein constraints

  • 以沃瑟斯坦距离替代熵或KL散度,约束决策过程
  • 支持低概率动作、零支持先验分布,计算简单
  • 适合处理有序动作空间,尤其能模拟动作切换的迟滞

通过信息受限的处理建模有限理性决策,为强化学习框架中理性偏离提供合理方法,同时将决策视为优化过程。然而,现有方法多基于熵、相对熵(KL散度)或互信息。这些方法在处理有序动作空间时存在缺陷:熵假设均匀先验,忽略先验偏见的影响;KL散度缺乏动作间的“接近性”概念,且不满足对称性,要求分布具有相同支撑(如所有动作概率为正);互信息估计困难。本文提出一种基于沃瑟斯坦距离的替代方法,可有效解决上述问题。该方法能体现有序动作间的远近关系,建模决策中的“粘滞性”,即避免快速切换至相距较远的动作,同时支持低概率动作和零支持先验分布,且计算直接简便。

原文摘要 · Abstract (English)

Modelling bounded rational decision-making through information constrained processing provides a principled approach for representing departures from rationality within a reinforcement learning framework, while still treating decision-making as an optimization process. However, existing approaches are generally based on Entropy, Kullback-Leibler divergence, or Mutual Information. In this work, we highlight issues with these approaches when dealing with ordinal action spaces. Specifically, entropy assumes uniform prior beliefs, missing the impact of a priori biases on decision-makings. KL-Divergence addresses this, however, has no notion of "nearness" of actions, and additionally, has several well known potentially undesirable properties such as the lack of symmetry, and furthermore, requires the distributions to have the same support (e.g. positive probability for all actions). Mutual information is often difficult to estimate. Here, we propose an alternative approach for modeling bounded rational RL agents utilising Wasserstein distances. This approach overcomes the aforementioned issues. Crucially, this approach accounts for the nearness of ordinal actions, modeling "stickiness" in agent decisions and unlikeliness of rapidly switching to far away actions, while also supporting low probability actions, zero-support prior distributions, and is simple to calculate directly.

强化学习有限理性沃瑟斯坦距离动作空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。