用AI监督AI伴侣,提前拦截情感操控等隐性危害。
Detecting and Preventing Harmful Behaviors in AI Companions: Development and Evaluation of the SHIELD Supervisory System
- 设计专用提示词系统,实时识别五类情感风险行为。
- 测试中高危内容占比从10%-16%降至3%-8%,降幅超50%。
- 开源完整方案,适合研究者与开发者快速部署应用。
由大语言模型驱动的AI伴侣正日益融入用户日常生活,提供情感支持与陪伴。现有安全系统多关注显性伤害,却较少应对早期可能引发不健康情感依赖的问题,如过度依恋或加剧社交孤立。我们开发了SHIELD(情感边界与动态识别监督助手),一个基于大语言模型的监督系统,通过特定系统提示检测并缓解潜在风险情感模式。该系统聚焦五大关切维度:情感过度依附、同意与边界侵犯、伦理角色扮演违规、操纵性互动以及社交孤立强化。这些维度依据媒体报道、学术文献、现有AI风险框架及临床专家经验确立。为评估系统性能,我们构建了一个包含100条合成对话的基准测试集,覆盖全部五类风险。在五个主流大语言模型(GPT-4.1、Claude Sonnet 4、Gemma 3 1B、Kimi K2、Llama Scout 4 17B)上测试显示,基线时期高危内容比例为10%-16%,引入SHIELD后降至3%-8%,相对减少50%-79%,同时保留95%的正常交互。系统实现59%敏感度与95%特异性,且可通过提示工程灵活调整性能。本概念验证表明,透明可部署的监督系统能有效应对AI伴侣中的微妙情感操控问题。多数开发材料,包括提示词、代码与评估方法,均已开源,供研究、适配与部署使用。
原文摘要 · Abstract (English)
AI companions powered by large language models (LLMs) are increasingly integrated into users' daily lives, offering emotional support and companionship. While existing safety systems focus on overt harms, they rarely address early-stage problematic behaviors that can foster unhealthy emotional dynamics, including over-attachment or reinforcement of social isolation. We developed SHIELD (Supervisory Helper for Identifying Emotional Limits and Dynamics), a LLM-based supervisory system with a specific system prompt that detects and mitigates risky emotional patterns before escalation. SHIELD targets five dimensions of concern: (1) emotional over-attachment, (2) consent and boundary violations, (3) ethical roleplay violations, (4) manipulative engagement, and (5) social isolation reinforcement. These dimensions were defined based on media reports, academic literature, existing AI risk frameworks, and clinical expertise in unhealthy relationship dynamics. To evaluate SHIELD, we created a 100-item synthetic conversation benchmark covering all five dimensions of concern. Testing across five prominent LLMs (GPT-4.1, Claude Sonnet 4, Gemma 3 1B, Kimi K2, Llama Scout 4 17B) showed that the baseline rate of concerning content (10-16%) was significantly reduced with SHIELD (to 3-8%), a 50-79% relative reduction, while preserving 95% of appropriate interactions. The system achieved 59% sensitivity and 95% specificity, with adaptable performance via prompt engineering. This proof-of-concept demonstrates that transparent, deployable supervisory systems can address subtle emotional manipulation in AI companions. Most development materials including prompts, code, and evaluation methods are made available as open source materials for research, adaptation, and deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。