arXiv:2502.19320cs.CLcs.AI2025-02ICLR被引 8

给大模型加安全罩,确保它不会说跑题的话。

Shh, don't say that! Domain Certification in LLMs

  • 提出领域认证机制,量化大模型跑偏风险。
  • 新方法VALID能紧致绑定越界输出概率,拒答率低。
  • 适合客服、医疗等需严格控域的场景使用。

大型语言模型常被用于特定任务,如客户服务机器人,依赖其广泛的语言理解能力提升表现。然而,这些模型易受对抗攻击,可能生成非预期领域的输出。为形式化、评估并缓解此风险,本文引入领域认证——一种准确刻画语言模型域外行为的保证。随后提出简单有效的方法VALID,可提供对抗边界作为证书。在多种数据集上的实验表明,该方法能生成有意义的证书,在最小化拒答行为代价下,紧密约束域外样本的概率。

原文摘要 · Abstract (English)

Large language models (LLMs) are often deployed to perform constrained tasks, with narrow domains. For example, customer support bots can be built on top of LLMs, relying on their broad language understanding and capabilities to enhance performance. However, these LLMs are adversarially susceptible, potentially generating outputs outside the intended domain. To formalize, assess, and mitigate this risk, we introduce domain certification; a guarantee that accurately characterizes the out-of-domain behavior of language models. We then propose a simple yet effective approach, which we call VALID that provides adversarial bounds as a certificate. Finally, we evaluate our method across a diverse set of datasets, demonstrating that it yields meaningful certificates, which bound the probability of out-of-domain samples tightly with minimum penalty to refusal behavior.

大模型安全领域认证对抗防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。