arXiv:2410.04683cs.LGcs.AI2024-10被引 6

提出可计算的目标导向性度量方法,评估AI是否隐含追求未知目标。

Towards Measuring Goal-Directedness in AI Systems

  • 基于策略在稀疏奖励下的近最优表现,定义目标导向性
  • 在简单马尔可夫决策过程环境中验证该度量的有效性
  • 适用于大模型等前沿AI系统,为安全评估提供新思路

深度学习的进展引发了对通用人工智能系统超越人类能力的期待,但若这些系统追求非预期目标,可能带来灾难性后果。关键问题在于:它们是否具备连贯且目标导向的行为模式,即优化某个未知目标?尽管已有研究尝试评估此类行为,但现有最严格的目标导向性定义在真实场景中难以计算。本文基于前人工作,探索强化学习环境中的策略目标导向性,提出一种新定义:衡量策略是否能被建模为多个稀疏奖励函数下的近最优解。我们对该定义进行了操作化,并在玩具马尔可夫决策过程(MDP)环境中测试。此外,还探讨了其在前沿大语言模型(LLMs)中的应用潜力。本研究贡献在于提出一种更简单、更易计算的目标导向性定义,有助于判断AI系统是否存在危险目标追求倾向。建议进一步探索基于此框架的连贯性与目标导向性测量方法。

原文摘要 · Abstract (English)

Recent advances in deep learning have brought attention to the possibility of creating advanced, general AI systems that outperform humans across many tasks. However, if these systems pursue unintended goals, there could be catastrophic consequences. A key prerequisite for AI systems pursuing unintended goals is whether they will behave in a coherent and goal-directed manner in the first place, optimizing for some unknown goal; there exists significant research trying to evaluate systems for said behaviors. However, the most rigorous definitions of goal-directedness we currently have are difficult to compute in real-world settings. Drawing upon this previous literature, we explore policy goal-directedness within reinforcement learning (RL) environments. In our findings, we propose a different family of definitions of the goal-directedness of a policy that analyze whether it is well-modeled as near-optimal for many (sparse) reward functions. We operationalize this preliminary definition of goal-directedness and test it in toy Markov decision process (MDP) environments. Furthermore, we explore how goal-directedness could be measured in frontier large-language models (LLMs). Our contribution is a definition of goal-directedness that is simpler and more easily computable in order to approach the question of whether AI systems could pursue dangerous goals. We recommend further exploration of measuring coherence and goal-directedness, based on our findings.

目标导向性强化学习大模型安全智能评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。