arXiv:2506.13593cs.LGstat.AP2025-06被引 3

提出新安全指标,量化大模型生成有害内容所需次数

Calibrated Predictive Lower Bounds on Time-to-Unsafe-Sampling in LLMs

  • 用生存分析与置信推断,构建生成有害内容时间的下界预测
  • 在真实与合成数据上验证,可有效评估提示词的安全风险
  • 适合关注生成式AI安全评估的研究者与开发者

我们提出一种新的生成模型安全度量——时间到不安全采样(time-to-unsafe-sampling),定义为大型语言模型(LLM)在生成过程中首次出现不安全(如有毒)输出所需的生成次数。该指标为提示词自适应安全评估提供了新维度。然而,由于安全模型中不安全输出通常极稀少,在合理采样预算下可能无法观测到。为此,我们将估计问题建模为生存分析,并基于最近的拟合预测理论,提出一种新颖的校准技术,以构建给定提示词下时间到不安全采样的严格覆盖保证的下界预测(LPB)。关键技术创新在于优化的采样预算分配方案,在保持无分布假设保证的前提下显著提升样本效率。在合成与真实数据上的实验支持了理论结果,并展示了该方法在生成式AI安全风险评估中的实际价值。

原文摘要 · Abstract (English)

We introduce time-to-unsafe-sampling, a novel safety measure for generative models, defined as the number of generations required by a large language model (LLM) to trigger an unsafe (e.g., toxic) response. While providing a new dimension for prompt-adaptive safety evaluation, quantifying time-to-unsafe-sampling is challenging: unsafe outputs are often rare in well-aligned models and thus may not be observed under any feasible sampling budget. To address this challenge, we frame this estimation problem as one of survival analysis. We build on recent developments in conformal prediction and propose a novel calibration technique to construct a lower predictive bound (LPB) on the time-to-unsafe-sampling of a given prompt with rigorous coverage guarantees. Our key technical innovation is an optimized sampling-budget allocation scheme that improves sample efficiency while maintaining distribution-free guarantees. Experiments on both synthetic and real data support our theoretical results and demonstrate the practical utility of our method for safety risk assessment in generative AI models.

安全评估语言模型生存分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。