arXiv:2607.18454cs.LGcs.AI2026-07

提出新方法精准评估语言模型罕见故障风险,解决传统方法易失效问题。

Estimating Rare Events in Language Models with Proper Evaluation

论文配图:Estimating Rare Events in Language Models with Proper Evaluation
图 1 · 摘自论文原文
  • 基于梯度的马尔可夫链蒙特卡洛,在连续激活空间中搜索罕见事件
  • 在小模型上显著降低对数误差,最优时比最强基线提升30%以上
  • 适合高风险场景下的安全评估,尤其关注低估与高估代价不均的情况

量化语言模型在对抗性分布偏移或大规模部署下罕见故障的风险,需估计极低概率事件,常规随机采样无法胜任。现有方法在极端稀有情形下易出现零估计崩溃或系统性偏差,标准评估损失也常不稳定或不匹配不对称安全成本。本文提出梯度激活自适应多层分裂(GA-AMLS),将罕见事件蒙特卡洛方法拓展至语言模型的连续激活空间。该方法采用基于梯度的MCMC核导航激活空间,避免输入空间搜索中的零估计崩溃,并以显式重尾激活先验替代以往方法的独立假设。同时提出漂移幂Bregman(SPB)损失,一种有限且可调不对称惩罚的严格评分规则。小规模Transformer模型实验显示:在对称评估下GA-AMLS表现最佳,平均对数空间平方误差较最强基线更低;而在不对称惩罚下,存在高估偏差的方法更优。结果表明,估算器选择应与部署情境匹配。本研究确立激活空间为语言模型罕见事件估计的可行域,克服离散输入空间搜索的脆弱性。

原文摘要 · Abstract (English)

Quantifying the risk of rare failures in language models, such as those triggered by adversarial distribution shifts or very large-scale deployments, requires estimating probabilities far too small for random sampling. While recent work has formalized Low Probability Estimation, existing pipelines remain fragile in the rarest regimes: estimators can suffer zero-estimate collapse or systematic bias, and standard evaluation losses can become unstable or poorly matched to asymmetric safety costs. In this work, we introduce Gradient Activation Adaptive Multi-Level Splitting (GA-AMLS), which adapts rare-event Monte Carlo methods to the continuous activation space of language models. Specifically, GA-AMLS uses a gradient-based MCMC kernel to navigate activation space, eliminating the zero-estimate collapse of input-space search and replacing the independence assumptions of prior activation-space estimators with conditional sampling under an explicit, heavier-tailed activation prior. We also propose the Shifted-Power Bregman (SPB) Loss, a proper scoring rule that remains finite for zero-estimates and offers tunable asymmetry between underestimation and overestimation penalties. Experiments on small transformer models reveal a bias-variance tradeoff: GA-AMLS achieves the lowest loss under symmetric evaluation, reducing average log-space squared error relative to the strongest baseline across model sizes, while methods with overestimation bias prevail under asymmetric penalties. Our findings highlight that estimator choice should be matched to deployment context. More broadly, our work establishes activation space as a tractable domain for rare-event estimation in language models, circumventing the brittleness of discrete input-space search.

语言模型罕见事件风险评估安全评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。