arXiv:2501.05455cs.CYcs.AI2025-01被引 5

打通上游与下游AI安全,让通用模型安全与具体应用安全协同进化。

Upstream and Downstream AI Safety: Both on the Same River?

  • 区分通用模型风险(上游)与具体场景风险(下游)
  • 发现共性问题如系统失效模式可跨框架复用
  • 适合关注模型对齐与工程安全融合的研究者

传统安全工程关注系统在实际使用中的表现,如自动驾驶车辆的运行设计域(道路布局、限速、天气等),这称为下游安全。相比之下,前沿AI安全研究关注超越特定应用场景的因素,例如大语言模型逃避人类控制或生成有害内容(如制造炸弹)的能力,这称为上游安全。本文梳理了上下游安全框架的特征,并探讨两者之间潜在的协同效应。例如,下游安全中的共因失效概念能否用于评估AI防护机制的有效性?前沿AI的能力与局限是否可为下游安全分析提供支持,例如在微调大模型以计算自主船舶航行计划时?论文识别出若干有前景的研究方向,并指出实现两类安全框架融合所面临的挑战。

原文摘要 · Abstract (English)

Traditional safety engineering assesses systems in their context of use, e.g. the operational design domain (road layout, speed limits, weather, etc.) for self-driving vehicles (including those using AI). We refer to this as downstream safety. In contrast, work on safety of frontier AI, e.g. large language models which can be further trained for downstream tasks, typically considers factors that are beyond specific application contexts, such as the ability of the model to evade human control, or to produce harmful content, e.g. how to make bombs. We refer to this as upstream safety. We outline the characteristics of both upstream and downstream safety frameworks then explore the extent to which the broad AI safety community can benefit from synergies between these frameworks. For example, can concepts such as common mode failures from downstream safety be used to help assess the strength of AI guardrails? Further, can the understanding of the capabilities and limitations of frontier AI be used to inform downstream safety analysis, e.g. where LLMs are fine-tuned to calculate voyage plans for autonomous vessels? The paper identifies some promising avenues to explore and outlines some challenges in achieving synergy, or a confluence, between upstream and downstream safety frameworks.

AI安全上游安全下游安全协同机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。