arXiv:2411.04462cs.AIcs.GT2024-11被引 2

用模拟信念修正因果决策,让CDT也能选最优策略

Can CDT rationalise the ex ante optimal policy via modified anthropics?

  • 假设自己可能在预测者的模拟中,调整信念以影响决策
  • 在特定条件下,修正后的CDT能实现最优事前政策
  • 适合研究决策理论与认知偏见交叉的学者

在新康姆悖论中,因果决策理论(CDT)推荐双盒选择,与证据决策理论(EDT)及事前最优策略相悖。但若认为自己可能处于预测者为决定是否向不透明盒子放一百万美元而运行的模拟中,则CDT可能推荐单盒选择,以促使预测者填充盒子。本文研究该思路的推广:考虑一般化的类新康姆问题,在合理自定位信念下,使CDT推荐与EDT式的事前最优策略一致。我们考察两种方法:一种建模世界在运行代理的模拟,另一种不依赖此类模型(称作‘广义广义三分法’,GGT)。对每种方法,我们刻画相应的CDT策略,并证明在某些条件下,这些策略包含事前最优策略。

原文摘要 · Abstract (English)

In Newcomb's problem, causal decision theory (CDT) recommends two-boxing and thus comes apart from evidential decision theory (EDT) and ex ante policy optimisation (which prescribe one-boxing). However, in Newcomb's problem, you should perhaps believe that with some probability you are in a simulation run by the predictor to determine whether to put a million dollars into the opaque box. If so, then causal decision theory might recommend one-boxing in order to cause the predictor to fill the opaque box. In this paper, we study generalisations of this approach. That is, we consider general Newcomblike problems and try to form reasonable self-locating beliefs under which CDT's recommendations align with an EDT-like notion of ex ante policy optimisation. We consider approaches in which we model the world as running simulations of the agent, and an approach not based on such models (which we call 'Generalised Generalised Thirding', or GGT). For each approach, we characterise the resulting CDT policies, and prove that under certain conditions, these include the ex ante optimal policies.

决策理论因果推理模拟假说

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。