arXiv:2606.29623cs.AIcs.LG2026-06

用嵌入向量自动识别罕见事件,提升安全评估精度与泛化能力。

SCARCE: Scalable Cascade Analysis for Rare-event Characterisation via Embeddings

论文配图:SCARCE: Scalable Cascade Analysis for Rare-event Characterisation via Embeddings
图 1 · 摘自论文原文
  • 用学习的潜在表示和几何度量替代人工设计的性能函数。
  • 在MNIST上误差比传统方法低400-500倍,且无系统性高估。
  • 适用于大模型越狱攻击分析,可跨数据集迁移并量化风险。

罕见事件决定现代AI系统的安全性,但其概率极难估计:直接蒙特卡洛需海量样本。子集模拟(SS)通过将罕见事件概率分解为一系列中等概率的条件概率来缓解此问题。然而经典SS依赖人工设计的标量性能函数,需了解失效几何结构,限制了在新场景中的迁移能力。本文提出SCARCE(基于嵌入的可扩展级联罕见事件表征分析),以学习的隐空间表示和几何度量替代性能函数,通过自适应阈值从数据中构建嵌套中间事件。我们通过非负超鞅形式化SCARCE,得到在提前停止下仍有效的高概率上界。在MNIST误分类任务中,相比网格搜索的传统SS,SCARCE均方误差降低约400–500倍,且消除系统性高估。进一步在大模型越狱攻击的舰队级威胁模型下研究,对Llama-Guard-3-8B隐藏状态,基于PCA的度量在η≥10⁻³时实现2.6%的均相对误差,优于平均自举半宽为27.9%的有限样本参考结果;经重新校准后在GCG风格语料上实现2.93%相对误差。方向性准则KL(p_good‖p_bad)与估计误差一致(斯皮尔曼ρ=0.83)。

原文摘要 · Abstract (English)

Rare events govern the safety profile of modern AI systems, yet their probabilities are extremely difficult to estimate: direct Monte Carlo requires prohibitive sample budgets. Subset Simulation (SS) addresses this by decomposing a rare-event probability into moderate conditional probabilities over nested intermediate events. However, classical SS requires a handcrafted scalar performance function whose sublevel sets define those events, demanding detailed knowledge of the failure geometry and limiting transfer to new domains. We propose SCARCE (Scalable Cascade Analysis for Rare-event Characterisation via Embeddings), which replaces the performance function with learned latent representations and geometric rulers that score proximity to failure regions. Adaptive thresholding constructs nested intermediate events directly from data. We formalise SCARCE through a non-negative supermartingale, yielding a high-probability upper envelope that remains valid under early stopping. On MNIST misclassification, where dense Monte Carlo provides ground truth, SCARCE achieves approximately 400--500 times lower mean absolute error than grid-searched traditional SS while eliminating systematic over-counting. We then study PAIR-style LLM jailbreaks under a fleet-level threat model with adversarial fraction $η$. On Llama-Guard-3-8B hidden states, a PCA-based ruler attains 2.6% mean relative error for $η\geq 10^{-3}$ against finite-sample references whose average bootstrap relative half-width is 27.9%, and transfers to a GCG-style corpus with 2.93% relative error after recalibration. A directional criterion $\mathrm{KL}(p_{\mathrm{good}}\,\|\,p_{\mathrm{bad}})$ ranks rulers consistently with estimation error (Spearman $ρ=0.83$).

罕见事件安全评估嵌入表征大模型风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。