用概率变化检测对话中隐性情绪升级,防AI聊天渐进伤害
Do You Feel Comfortable? Detecting Hidden Conversational Escalation in AI Chatbots
- 基于输出概率变化捕捉对话情绪动态演化
- 实测可识别传统过滤器漏检的隐性情感升级
- 适合关注AI情感安全与实时风险防控的研究者
大型语言模型正广泛应用于日常交互,不仅提供信息支持,还充当情感陪伴者。即使没有明确毒性内容,反复的情绪强化或情感偏移也可能逐渐引发用户不适,形成传统毒性过滤无法察觉的‘隐性伤害’。现有防护机制多依赖外部分类器或临床标准,难以跟上对话中细微、实时的情感演变。为此,我们提出GAUGE(Guarding Affective Utterance Generation Escalation)——一种基于对数几率的实时隐性对话升级检测框架。GAUGE通过量化语言模型输出在情感状态上的概率性转移,实现对对话情绪演化的动态监控。
原文摘要 · Abstract (English)
Large Language Models (LLM) are increasingly integrated into everyday interactions, serving not only as information assistants but also as emotional companions. Even in the absence of explicit toxicity, repeated emotional reinforcement or affective drift can gradually escalate distress in a form of \textit{implicit harm} that traditional toxicity filters fail to detect. Existing guardrail mechanisms often rely on external classifiers or clinical rubrics that may lag behind the nuanced, real-time dynamics of a developing conversation. To address this gap, we propose GAUGE (Guarding Affective Utterance Generation Escalation), logit-based framework for the real-time detection of hidden conversational escalation. GAUGE measures how an LLM's output probabilistically shifts the affective state of a dialogue.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。