在噪声环境下区分恶意行为与执行错误,提升合作决策准确性。
Intention Inference Under Execution Noise: Separating Aleatoric and Epistemic Uncertainty in Social Dilemmas
- 用部分可观测马尔可夫决策模型将意图设为隐状态,动作作为噪声观测。
- 发现对称噪声存在临界阈值,超过后合作会崩溃且与先验学习相关。
- 适合需判断对手真实意图的博弈场景,尤其在条件合作中优势明显。
在存在执行噪声的社交困境中,意图动作会被随机干扰,导致观察到的背叛可能源于敌意或执行错误。标准马尔可夫决策过程将执行动作视为状态,无法区分二者,引发系统性过度报复。本文提出一种部分可观测马尔可夫决策过程(POMDP)框架,将对手意图建模为隐状态,执行动作作为噪声观测,并在主动推理(AIF)框架下求解,其代价函数分解为认知与实用两部分,共同实现当前意图推断与意图演化学习。在对称噪声下的重复囚徒困境中,我们推导出一个决定合作崩溃的临界噪声阈值,该阈值与学习到的先验分布固定点相关。实验表明,意图推断的价值依赖于上下文:在条件合作对手面前,该方法持续占优;但在高噪声下双方同时进行意图推断时,会因信念关联导致合作集体崩溃。该优势仅存在于意图归因影响决策的游戏场景中。
原文摘要 · Abstract (English)
In noisy social dilemmas, intended actions are stochastically corrupted before execution, so an observed defection may reflect hostile intent or action error. Standard Markov Decision Process (MDP) formulations treat executed actions as states, structurally precluding this distinction and causing systematic over-retaliation. We introduce a Partially Observable MDP (POMDP) formulation encoding opponent intentions as latent states and executed actions as noisy observations, solved within the active inference (AIF) framework with a cost function that decomposes into epistemic and pragmatic components that jointly address inferring current intent and learning how intent evolves. In the Iterated Prisoner's Dilemma with symmetric noise, we derive a critical noise threshold governing cooperation collapse, connecting it to a fixed-point condition on learned priors. Experiments reveal that the value of intention inference is context-dependent: the POMDP provides consistent advantages against conditionally cooperative opponents, but mutual intention inference under sufficient noise produces correlated belief-driven collapse. The advantage is specific to games where intent attribution is decision-relevant.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。