arXiv:2510.03969cs.AIcs.CR2025-10中稿 · ICLR被引 1

为大模型对话安全提供统计保证,能发现高达70%的灾难性响应风险。

How Catastrophic is Your LLM? Certifying Risk in Conversation

  • 构建对话概率图模型,用马尔可夫过程模拟真实对话流。
  • 在前沿模型中发现灾难性响应概率下限最高达70%。
  • 适合关注大模型安全、评估对话风险的研究者与开发者。

大型语言模型(LLMs)在对话场景中可能产生严重危害公共安全的灾难性响应。现有评估方法常因依赖固定攻击提示序列、缺乏统计保障且无法扩展至多轮对话空间而难以揭示此类漏洞。本文提出C$^3$LLM,一种新型、基于原理的统计认证框架,用于量化多轮对话中灾难性风险,并在对话分布上提供统计保证。我们将多轮对话建模为查询序列上的概率分布,通过查询图上的马尔可夫过程表示,边编码语义相似性以捕捉真实对话流动;利用置信区间量化灾难性风险。定义了多种低成本实用分布——随机节点、图路径及带拒绝的自适应分布。实验表明,这些分布可在前沿模型中揭示显著的灾难性风险,最差模型的认证下限高达70%,凸显当前前沿大模型亟需改进安全训练策略。

原文摘要 · Abstract (English)

Large Language Models (LLMs) can produce catastrophic responses in conversational settings that pose serious risks to public safety and security. Existing evaluations often fail to fully reveal these vulnerabilities because they rely on fixed attack prompt sequences, lack statistical guarantees, and do not scale to the vast space of multi-turn conversations. In this work, we propose C$^3$LLM, a novel, principled statistical Certification framework for Catastrophic risks in multi-turn Conversation for LLMs that bounds the probability of an LLM generating catastrophic responses under multi-turn conversation distributions with statistical guarantees. We model multi-turn conversations as probability distributions over query sequences, represented by a Markov process on a query graph whose edges encode semantic similarity to capture realistic conversational flow, and quantify catastrophic risks using confidence intervals. We define several inexpensive and practical distributions--random node, graph path, and adaptive with rejection. Our results demonstrate that these distributions can reveal substantial catastrophic risks in frontier models, with certified lower bounds as high as 70% for the worst model, highlighting the urgent need for improved safety training strategies in frontier LLMs.

大模型安全对话风险统计认证风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。