优化大模型拒答策略,让拒绝更温和,提升用户体验。
Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences
- 用部分回应替代完全拒绝,避免用户反感
- 部分回应使负面感知降低50%以上
- 现有模型和奖励机制不擅长使用温和拒答
当前大模型无论用户是否有恶意意图,都会拒绝潜在有害请求,导致安全与用户体验的权衡。通过对480名参与者评估3,840个问答对的研究发现,拒答策略显著影响用户感知,而用户真实动机影响甚微。部分合规——提供一般信息但不包含可操作细节——成为最优策略,使负面感知比完全拒绝降低超过50%。我们进一步分析了9个前沿大模型的响应模式,并评估6个奖励模型对不同拒答策略的评分,发现模型极少自然采用部分合规,且奖励模型普遍低估其价值。本研究指出,有效的安全防护应聚焦于设计有温度的拒答,而非识别用户意图,为实现安全与持续用户参与的平衡提供新路径。
原文摘要 · Abstract (English)
Current LLMs are trained to refuse potentially harmful input queries regardless of whether users actually had harmful intents, causing a tradeoff between safety and user experience. Through a study of 480 participants evaluating 3,840 query-response pairs, we examine how different refusal strategies affect user perceptions across varying motivations. Our findings reveal that response strategy largely shapes user experience, while actual user motivation has negligible impact. Partial compliance -- providing general information without actionable details -- emerges as the optimal strategy, reducing negative user perceptions by over 50% to flat-out refusals. Complementing this, we analyze response patterns of 9 state-of-the-art LLMs and evaluate how 6 reward models score different refusal strategies, demonstrating that models rarely deploy partial compliance naturally and reward models currently undervalue it. This work demonstrates that effective guardrails require focusing on crafting thoughtful refusals rather than detecting intent, offering a path toward AI safety mechanisms that ensure both safety and sustained user engagement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。