arXiv:2412.04758cs.AIcs.LG2024-12NeurIPS被引 9

提出可计算的目标导向性度量MEG,用于评估智能体的意图明确程度。

Measuring Goal-Directedness

论文配图:Measuring Goal-Directedness
图 1 · 摘自论文原文
  • 基于最大因果熵框架,定义可计算的目标导向性度量
  • 支持已知效用函数、假设类或随机变量集上的度量
  • 适用于评估AI行为意图,对安全与哲学研究有启发

我们提出了最大熵目标导向性(MEG),一种在因果模型和马尔可夫决策过程中的目标导向性形式化度量,并给出了其计算算法。衡量目标导向性至关重要,因其是评估人工智能潜在危害的关键要素,也具有哲学意义,因目标导向性是代理行为的核心特征。MEG基于逆强化学习中最大因果熵框架的改编,可针对已知效用函数、效用函数假设类或一组随机变量进行度量。我们证明了MEG满足若干理想性质,并通过小规模实验展示了算法有效性。

原文摘要 · Abstract (English)

We define maximum entropy goal-directedness (MEG), a formal measure of goal-directedness in causal models and Markov decision processes, and give algorithms for computing it. Measuring goal-directedness is important, as it is a critical element of many concerns about harm from AI. It is also of philosophical interest, as goal-directedness is a key aspect of agency. MEG is based on an adaptation of the maximum causal entropy framework used in inverse reinforcement learning. It can measure goal-directedness with respect to a known utility function, a hypothesis class of utility functions, or a set of random variables. We prove that MEG satisfies several desiderata and demonstrate our algorithms with small-scale experiments.

目标导向性智能体评估强化学习形式化度量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。