用量子思想解决决策中矛盾信息带来的不确定性问题
Quantum-Inspired Reinforcement Learning in the Presence of Epistemic Ambivalence
- 引入量子态概念建模矛盾证据下的决策过程
- 在双状态和格子问题中实现最优策略收敛
- 适合研究复杂认知冲突的强化学习方向
在线决策中的不确定性源于已知策略利用与新可能性探索之间的平衡难题。本文聚焦一种特殊不确定性——认识性歧义(epistemic ambivalence, EA),它由相互冲突的证据或矛盾经验引发,导致不确定性和信心间的微妙互动,且不会随新知识增加而消减。为此,我们提出一种新型框架——认识性歧义马尔可夫决策过程(EA-MDP),借鉴量子力学中的量子态概念,评估每种可能结果的概率与回报。通过量子测量技术计算奖励函数,并证明了该框架下最优策略与最优价值函数的存在性。进一步提出EA-epsilon-greedy Q-learning算法。在双状态问题与格子问题两个实验场景中验证,结果表明该方法能在存在EA的情况下使智能体收敛至最优策略。
原文摘要 · Abstract (English)
The complexity of online decision-making under uncertainty stems from the requirement of finding a balance between exploiting known strategies and exploring new possibilities. Naturally, the uncertainty type plays a crucial role in developing decision-making strategies that manage complexity effectively. In this paper, we focus on a specific form of uncertainty known as epistemic ambivalence (EA), which emerges from conflicting pieces of evidence or contradictory experiences. It creates a delicate interplay between uncertainty and confidence, distinguishing it from epistemic uncertainty that typically diminishes with new information. Indeed, ambivalence can persist even after additional knowledge is acquired. To address this phenomenon, we propose a novel framework, called the epistemically ambivalent Markov decision process (EA-MDP), aiming to understand and control EA in decision-making processes. This framework incorporates the concept of a quantum state from the quantum mechanics formalism, and its core is to assess the probability and reward of every possible outcome. We calculate the reward function using quantum measurement techniques and prove the existence of an optimal policy and an optimal value function in the EA-MDP framework. We also propose the EA-epsilon-greedy Q-learning algorithm. To evaluate the impact of EA on decision-making and the expedience of our framework, we study two distinct experimental setups, namely the two-state problem and the lattice problem. Our results show that using our methods, the agent converges to the optimal policy in the presence of EA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。