提出渐进式干预框架,平衡AI情感陪伴的安全与关系信任。
SLIP & ETHICS: Graduated Intervention for AI Emotional Companions
- 基于情绪强度与叙事动态设计三阶干预机制
- 零误报检测危机人格,高能持续8天无干预暴露边界
- 模型能力提升可增强安全检测,适合情感计算研究者
AI情感陪伴面临安全-信任悖论:严格防护会破坏支持性关系,放任则可能危害用户。本文提出四阶段渐进干预协议SLIP,依据情绪强度(a)和叙事动态(m)生成无、软、硬三级干预措施,并引入ETHICS(人类-AI互动情境信号的涌现分类法),强调“信号而非标签”。结合小规模生产部署(N=68条记录,10名用户,10周)与合成人格测试(N=91,5种行为风险类型),流型人格实现0%假阳性,危机人格显示预期升级模式。但初始结果发现,连续8天高能量状态未触发任何干预(0/8),暴露出“不病理化”原则与安全需求的冲突边界。后续三模型压力测试表明,模型能力提升使检测率从0/8升至6/8,同时保持最大模型0/10的流型假阳性,验证渐进干预是缓解安全-信任张力的设计方向。
原文摘要 · Abstract (English)
AI emotional companions face a safety-rapport paradox: restrictive safeguards can damage supportive alliance, while permissive systems risk user harm. We present SLIP (Staged Layers of Intervention Protocol), a four-stage graduated methodology deriving interventions (none, soft, hard) from structured qualitative indicators -- affect intensity (a) and narrative dynamism (m) -- alongside ETHICS (Emergent Taxonomy for Human-AI Interaction Context Signals), a "signals not labels" taxonomy. An evaluation combining a small-scale production deployment (N=68 entries, 10 users, 10 weeks) with a synthetic persona battery (N=91, 5 behavioral-risk profiles) achieved 0% false positives for the flow persona and showed expected escalation patterns in crisis-oriented personas. However, initial results showed that 8 consecutive days of high-energy elevation produced zero interventions (0/8), exposing a boundary where the "do not pathologize" principle conflicts with safety. A subsequent three-model stress test demonstrated that increased model capability improves detection from 0/8 to 6/8 while preserving 0/10 flow false positives in the largest model. Read as preliminary, these findings position graduated intervention as a design direction for navigating -- not resolving -- the safety-rapport tension in affective computing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。