arXiv:2602.05164cs.LGcs.AI2026-02被引 3

模型能力管控应独立于对齐目标,分层设防防止滥用。

Position: Capability Control Should be a Separate Goal From Alignment

  • 从数据、学习、系统三层面分层控制模型行为
  • 不同层级单独使用易失效,需组合防御
  • 适合关注安全可控的模型部署者

基础模型在广泛数据上训练,具备通用能力,但也扩大了误用和失效的可能性。本文主张将能力控制——对模型行为施加限制——视为独立于对齐的目标。对齐通常依赖上下文与偏好,而能力控制旨在设定硬性行为边界,包括对抗性诱导下的限制。我们提出将能力控制机制按模型生命周期分为三层:(i) 基于数据的训练分布控制,(ii) 基于权重或表征的训练中干预,(iii) 部署后对输入、输出和动作的系统级防护。由于各层单独使用存在固有缺陷,建议采用纵深防御策略,整合多层互补控制。此外,文章指出实现该控制的关键挑战,包括知识的双重用途性与组合泛化问题。

原文摘要 · Abstract (English)

Foundation models are trained on broad data distributions, yielding generalist capabilities that enable many downstream applications but also expand the space of potential misuse and failures. This position paper argues that capability control -- imposing restrictions on permissible model behavior -- should be treated as a distinct goal from alignment. While alignment is often context and preference-driven, capability control aims to impose hard operational limits on permissible behaviors, including under adversarial elicitation. We organize capability control mechanisms across the model lifecycle into three layers: (i) data-based control of the training distribution, (ii) learning-based control via weight- or representation-level interventions, and (iii) system-based control via post-deployment guardrails over inputs, outputs, and actions. Because each layer has characteristic failure modes when used in isolation, we advocate for a defense-in-depth approach that composes complementary controls across the full stack. We further outline key open challenges in achieving such control, including the dual-use nature of knowledge and compositional generalization.

模型安全能力控制防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。