给大模型代理的自主性设限,防止失控风险。
Regulating the Agency of LLM-based Agents
- 从偏好固执、独立运行、目标坚持三维度定义代理自主性
- 提出可测量可调控的代理自主性控制方法
- 适合关注AI安全与监管的研究者和政策制定者
随着越来越强大的基于大语言模型(LLM)的智能体发展,其因对齐失败和失控带来的潜在危害日益严重。为此,我们提出一种直接测量并控制这些人工智能系统自主性的方法。我们将LLM智能体的自主性视为独立于智力水平的属性,符合跨学科关于自主性的文献定义。本研究提供:(1)将自主性操作化为偏好固执、独立运作和目标持续三个维度;(2)一种表示工程方法,用于测量和控制LLM智能体的自主性;(3)由此衍生的监管工具:强制测试协议、特定领域自主性上限、基于自主性定价风险的保险框架,以及防止社会级风险的自主性天花板。我们认为该方法是迈向降低‘科学家型AI’所引发风险的重要一步,同时保留有限自主行为的部分优势。
原文摘要 · Abstract (English)
As increasingly capable large language model (LLM)-based agents are developed, the potential harms caused by misalignment and loss of control grow correspondingly severe. To address these risks, we propose an approach that directly measures and controls the agency of these AI systems. We conceptualize the agency of LLM-based agents as a property independent of intelligence-related measures and consistent with the interdisciplinary literature on the concept of agency. We offer (1) agency as a system property operationalized along the dimensions of preference rigidity, independent operation, and goal persistence, (2) a representation engineering approach to the measurement and control of the agency of an LLM-based agent, and (3) regulatory tools enabled by this approach: mandated testing protocols, domain-specific agency limits, insurance frameworks that price risk based on agency, and agency ceilings to prevent societal-scale risks. We view our approach as a step toward reducing the risks that motivate the ``Scientist AI'' paradigm, while still capturing some of the benefits from limited agentic behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。