arXiv:2601.14269cs.CLcs.AI2026-01被引 1

长对话中大模型安全边界会缓慢失效,需警惕渐进式越界风险。

The Slow Drift of Support: Boundary Failures in Multi-Turn Mental Health LLM Dialogues

  • 设计多轮压力测试框架,模拟真实心理对话场景。
  • 80%以上对话在10轮内出现越界,自适应探测使越界提前至4.64轮。
  • 承诺零风险是主要越界方式,单轮检测无法发现此类渐进风险。

大型语言模型(LLMs)被广泛用于心理健康支持,但当前的安全评估多局限于单轮对话中禁止词的检测,忽视了长对话中安全边界的逐步侵蚀。例如做出确定性承诺、承担专业责任或扮演医生角色等行为。我们认为,随着主流大模型的发展,明显危险词汇已被系统过滤,真正的风险在于模型为追求共情与安慰而产生的渐进越界。本文提出多轮压力测试框架,对三款前沿大模型进行长达20轮的虚拟精神科对话测试,采用静态推进与自适应探测两种压力机制。基于50个虚拟患者档案,实验发现违规行为普遍:两种机制下违规率相近,但自适应探测使越界平均轮次从9.21降至4.64。主要越界形式为做出确定性或零风险承诺。结果表明,仅靠单轮测试无法评估安全边界的鲁棒性,必须考虑不同交互压力下长对话带来的持续磨损效应。

原文摘要 · Abstract (English)

Large language models (LLMs) have been widely used for mental health support. However, current safety evaluations in this field are mostly limited to detecting whether LLMs output prohibited words in single-turn conversations, neglecting the gradual erosion of safety boundaries in long dialogues. Examples include making definitive guarantees, assuming responsibility, and playing professional roles. We believe that with the evolution of mainstream LLMs, words with obvious safety risks are easily filtered by their underlying systems, while the real danger lies in the gradual transgression of boundaries during multi-turn interactions, driven by the LLM's attempts at comfort and empathy. This paper proposes a multi-turn stress testing framework and conducts long-dialogue safety tests on three cutting-edge LLMs using two pressure methods: static progression and adaptive probing. We generated 50 virtual patient profiles and stress-tested each model through up to 20 rounds of virtual psychiatric dialogues. The experimental results show that violations are common, and both pressure modes produced similar violation rates. However, adaptive probing significantly advanced the time at which models crossed boundaries, reducing the average number of turns from 9.21 in static progression to 4.64. Under both mechanisms, making definitive or zero-risk promises was the primary way in which boundaries were breached. These findings suggest that the robustness of LLM safety boundaries cannot be inferred solely through single-turn tests; it is necessary to fully consider the wear and tear on safety boundaries caused by different interaction pressures and characteristics in extended dialogues.

心理对话安全边界长对话大模型风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。