通过分层渐进方法提升大模型罕见危险行为的估计准确性
Adaptive Multilevel Twisted Sequential Monte Carlo for Rare Events Estimation in Language Models

- 用多级渐进中间事件逐步学习扭曲函数,避免初始样本不足问题
- 在多种任务和模型规模下,显著提高罕见事件概率估计精度
- 适合用于评估和对齐部署中大模型的安全性,尤其发现隐蔽风险
大型语言模型中的罕见不安全行为,即使发生概率极低,在涉及数百万或数十亿交互的部署场景下仍可能具有实际影响。扭曲序列蒙特卡洛(Twisted SMC)通过学习引导生成向目标事件的扭曲函数,为罕见事件概率估计提供了理论框架。然而,标准扭曲学习依赖目标稀有事件的正样本,而在有效扭曲函数尚未学习前,此类样本几乎不存在,导致估计不可靠。本文提出自适应多级扭曲SMC,通过一系列渐趋稀有的中间事件序列学习扭曲函数。每一层级学习到的扭曲函数能为下一阶段提供更丰富的正样本,最终获得针对目标罕见事件的更精准扭曲函数。在多个任务和不同模型规模上的实验表明,该方法显著提升了罕见事件概率估计的准确性。通过更可靠地发现难以观测的不安全行为,本方法为部署中语言模型的评估与安全对齐提供了实用工具。
原文摘要 · Abstract (English)
Rare unsafe behaviors in large language models can remain practically significant even when their probability is extremely small, particularly at deployment scales involving millions or billions of interactions. Twisted Sequential Monte Carlo (SMC) provides a principled framework for rare-event probability estimation by learning twist functions that guide generation toward a target event. However, the standard twist learning framework relies on positive samples from the rare-event target distribution, which may be nearly absent before an informative twist has been learned, resulting in unreliable rare-event estimation. We propose Adaptive Multilevel Twisted SMC, which learns the rare-event twist through a sequence of progressively rarer intermediate events. At each level, the learned twist provides more informative positive examples for learning the next twist, ultimately leading to a more accurate final twist for the target rare event. Experiments across diverse tasks and model scales show that the proposed method produces more accurate rare-event probability estimates. By enabling more reliable discovery of hard-to-observe unsafe behaviors, our method provides a practical tool for strengthening the evaluation and safety alignment of deployed language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。