arXiv:2602.06791cs.LGcond-mat.dis-nn2026-02被引 3

系统分析大模型中罕见但重要的异常行为。

Rare Event Analysis of Large Language Models

  • 构建端到端框架,系统识别大模型的罕见行为。
  • 通过高效生成与概率估计,发现未在训练中出现的极端案例。
  • 适用于多种模型和场景,适合关注模型安全与可靠性的研究者。

作为概率模型,大语言模型(LLMs)在推理过程中会表现出罕见事件:即与典型行为相去甚远但极具重要性的行为。由于定义上罕见事件难以观测,而大模型部署规模巨大,开发阶段未见的事件在实际使用中可能变得显著。本文提出一个端到端的罕见事件分析框架,涵盖理论、高效生成策略、概率估计与误差分析,并通过具体实例加以验证。我们还拓展了该方法在其他模型与场景中的应用,凸显其概念与技术的普适性。

原文摘要 · Abstract (English)

Being probabilistic models, during inference large language models (LLMs) display rare events: behaviour that is far from typical but highly significant. By definition all rare events are hard to see, but the enormous scale of LLM usage means that events completely unobserved during development are likely to become prominent in deployment. Here we present an end-to-end framework for the systematic analysis of rare events in LLMs. We provide a practical implementation spanning theory, efficient generation strategies, probability estimation and error analysis, which we illustrate with concrete examples. We outline extensions and applications to other models and contexts, highlighting the generality of the concepts and techniques presented here.

大模型分析罕见事件推理安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。