arXiv:2409.18203cs.HCcs.AI2024-09被引 11

用地图导航式方法,精准管控大模型的复杂行为边界。

Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors

  • 借鉴制图思路,聚焦关键行为区域,抽象无关细节。
  • 通过交互工具定义规则,对暴力等敏感内容自动重写。
  • 适合安全专家设计模型行为规范,提升可控性。

AI政策需界定大语言模型(LLMs)的可接受行为范围,但面对庞大的行为空间,全面覆盖极具挑战。本文提出「政策地图」(Policy Maps),受物理地图绘制启发,不追求全量覆盖,而是通过有意识的设计选择,明确捕捉哪些行为特征、抽象哪些细节,实现高效导航。结合「政策投影仪」(Policy Projector)这一交互工具,开发者可探索模型输入输出的分布,自定义行为区域(如“暴力”),并以条件规则(如若输出含“暴力”且“具体描述”,则重写为无细节版本)干预模型输出。该工具支持基于LLM分类与引导的交互式策略编写,并可视化呈现设计过程。12位AI安全专家参与评估表明,系统能有效帮助制定针对错误性别假设、即时身体安全威胁等高风险行为的政策。

原文摘要 · Abstract (English)

AI policy sets boundaries on acceptable behavior for AI models, but this is challenging in the context of large language models (LLMs): how do you ensure coverage over a vast behavior space? We introduce policy maps, an approach to AI policy design inspired by the practice of physical mapmaking. Instead of aiming for full coverage, policy maps aid effective navigation through intentional design choices about which aspects to capture and which to abstract away. With Policy Projector, an interactive tool for designing LLM policy maps, an AI practitioner can survey the landscape of model input-output pairs, define custom regions (e.g., "violence"), and navigate these regions with if-then policy rules that can act on LLM outputs (e.g., if output contains "violence" and "graphic details," then rewrite without "graphic details"). Policy Projector supports interactive policy authoring using LLM classification and steering and a map visualization reflecting the AI practitioner's work. In an evaluation with 12 AI safety experts, our system helps policy designers craft policies around problematic model behaviors such as incorrect gender assumptions and handling of immediate physical safety threats.

AI安全行为控制提示工程大模型治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。