构建新康姆类决策问题数据集,测试大模型的推理与立场差异。
A dataset of questions on decision-theoretic reasoning in Newcomb-like problems
- 设计自然语言问题集,涵盖可验证答案与学界争议问题。
- 发现高推理能力模型更倾向支持因果决策论,且立场一致。
- 适用于评估大模型在复杂协作场景中的决策逻辑与态度。
我们引入了一个关于新康姆类决策问题的自然语言问题数据集。新康姆类问题包括代理与相似代理互动的情形,需推理对方可能采用类似思维模式。评估大模型在此类问题上的推理能力至关重要,因基于基础模型的代理间交互常具有新康姆类特征。某些推理方式可能促进模型间的合作。该数据集包含能力型问题(有唯一正确答案)与态度型问题(涉及决策理论学家间的分歧)。我们利用该数据集研究现有模型(OpenAI、Anthropic、Meta、GDM、Reka等)的决策能力与态度表现及其相互关系,以及简单提示干预下的变化。结果表明:不同模型的态度差异显著;高能力模型更倾向于支持证据决策论;且态度在不同类型问题中保持一致。
原文摘要 · Abstract (English)
We introduce a dataset of natural-language questions in the decision theory of so-called Newcomb-like problems. Newcomb-like problems include, for instance, decision problems in which an agent interacts with a similar other agent, and thus has to reason about the fact that the other agent will likely reason in similar ways. Evaluating LLM reasoning about Newcomb-like problems is important because interactions between foundation-model-based agents will often be Newcomb-like. Some ways of reasoning about Newcomb-like problems may allow for greater cooperation between models. Our dataset contains both capabilities questions (i.e., questions with a unique, uncontroversially correct answer) and attitude questions (i.e., questions about which decision theorists would disagree). We use our dataset for an investigation of decision-theoretical capabilities and expressed attitudes and their interplay in existing models (different models by OpenAI, Anthropic, Meta, GDM, Reka, etc.), as well as models under simple prompt-based interventions. We find, among other things, that attitudes vary significantly between existing models; that high capabilities are associated with attitudes more favorable toward so-called evidential decision theory; and that attitudes are consistent across different types of questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。