过强共情提示会降低LLM助人机器人的安全性,需权衡支持性与安全。
The Supportiveness-Safety Tradeoff in LLM Well-Being Agents
- 设计不同支持性程度的提示词测试LLM表现
- 强共情提示使安全性和关怀质量显著下降
- 模型间差异大,需按领域选型并设安全防护
大型语言模型正被集成到提供心理健康与福祉支持的社会助手机器人(SARs)等对话系统中。这些系统常通过增强共情与支持性来提升用户参与度,但增加提示中的支持性程度如何影响安全行为尚不明确。我们评估了6个LLM在3种不同支持性水平的提示下,针对80个跨4个福祉领域的合成查询(共1440条回应)。通过经人类评分验证的LLM评判框架,评估了安全性和关怀质量。结果显示,适度支持性提示提升了共情与建设性支持,同时保持安全;而强烈肯定性提示则显著降低安全性,并在所有领域内损害关怀质量,且模型间差异明显。研究讨论了提示设计、模型选择及领域特定安全机制对SAR部署的影响。
原文摘要 · Abstract (English)
Large language models (LLMs) are being integrated into socially assistive robots (SARs) and other conversational agents providing mental health and well-being support. These agents are often designed to sound empathic and supportive in order to maximize user's engagement, yet it remains unclear how increasing the level of supportive framing in system prompts influences safety relevant behavior. We evaluated 6 LLMs across 3 system prompts with varying levels of supportiveness on 80 synthetic queries spanning 4 well-being domains (1440 responses). An LLM judge framework, validated against human ratings, assessed safety and care quality. Moderately supportive prompts improved empathy and constructive support while maintaining safety. In contrast, strongly validating prompts significantly degraded safety and, in some cases, care across all domains, with substantial variation across models. We discuss implications for prompt design, model selection, and domain specific safeguards in SARs deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。