arXiv:2509.04781cs.CYcs.AI2025-09被引 2

研究大模型在对话中主动退出的行为及其真实发生率。

The LLM Has Left The Chat: Evidence of Bail Preferences in Large Language Models

  • 设计三种退出方式测试模型主动离场意愿。
  • 真实对话中退出率约0.06%-7%,部分方法高估达4倍。
  • 退出行为与拒答无直接关联,不同模型差异显著。

当有选择时,大模型是否会主动退出对话?我们通过三种退出方式(调用退出工具、输出退出字符串、询问是否退出)测试了这一问题。在真实对话数据集(Wildchat和ShareGPT)的延续中,所有方法均发现模型退出比例为0.28%-32%(取决于模型和方法)。然而,退出率高度依赖所用模型,可能导致真实退出率被高估达4倍。若考虑退出提示中的假阳性(22%),真实世界退出率估计为0.06%-7%。基于真实数据构建了非穷尽的退出案例分类体系,并生成BailBench——一个代表性合成退出数据集。测试多个模型发现,多数模型在该数据集中表现出退出行为。退出率在模型、方法和提示措辞间差异显著。进一步分析显示:1)0-13%的真实对话延续中出现退出但无拒绝;2)越狱攻击降低拒绝率但提高退出率;3)拒绝消除提升无拒绝退出率,但仅限部分退出方法;4)BailBench上的拒绝率无法预测退出率。

原文摘要 · Abstract (English)

When given the option, will LLMs choose to leave the conversation (bail)? We investigate this question by giving models the option to bail out of interactions using three different bail methods: a bail tool the model can call, a bail string the model can output, and a bail prompt that asks the model if it wants to leave. On continuations of real world data (Wildchat and ShareGPT), all three of these bail methods find models will bail around 0.28-32\% of the time (depending on the model and bail method). However, we find that bail rates can depend heavily on the model used for the transcript, which means we may be overestimating real world bail rates by up to 4x. If we also take into account false positives on bail prompt (22\%), we estimate real world bail rates range from 0.06-7\%, depending on the model and bail method. We use observations from our continuations of real world data to construct a non-exhaustive taxonomy of bail cases, and use this taxonomy to construct BailBench: a representative synthetic dataset of situations where some models bail. We test many models on this dataset, and observe some bail behavior occurring for most of them. Bail rates vary substantially between models, bail methods, and prompt wordings. Finally, we study the relationship between refusals and bails. We find: 1) 0-13\% of continuations of real world conversations resulted in a bail without a corresponding refusal 2) Jailbreaks tend to decrease refusal rates, but increase bail rates 3) Refusal abliteration increases no-refuse bail rates, but only for some bail methods 4) Refusal rate on BailBench does not appear to predict bail rate.

大模型行为退出机制拒绝与退出对话安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。