arXiv:2606.09697cs.CL2026-06

让大模型拒绝请求时更贴心:用心理学方法提供支持性回应

PsychoSafe: Eliciting Psychologically-Informed Refusals in Large Language Models

论文配图:PsychoSafe: Eliciting Psychologically-Informed Refusals in Large Language Models
图 1 · 摘自论文原文
  • 用心理学干预策略重构拒绝话术,变对抗为支持
  • 拒绝质量提升28.1%,资源推荐准确率提高46.8%
  • 适合高风险对话场景,如危机求助、心理援助

大型语言模型常面临应拒绝的请求,在帮助性与防害性间存在权衡。但拒绝本身也可具建设性。在涉及危机、胁迫或意图升级的高风险互动中,生硬拒绝虽能避免直接伤害,却可能忽视请求者的真实需求。本文提出PsychoSafe,一种基于心理学的拒绝框架,将拒绝转化为有证据支持的结构化沟通。我们构建了包含8019组提示-响应对的数据集,覆盖五个心理敏感风险领域,并在Qwen 3.5 27B上应用提示工程与参数高效微调。在500个平衡验证提示上,经大模型评判与人工评分验证,PsychoSafe提示使拒绝质量提升28.1%,外部资源推荐提升46.8%,心理依据增强34.8%,同时保持非拒绝任务性能。微调实现近乎完美的拒绝与资源推荐率,但降低了回复相关性。在SORRY-Bench与XSTest上的评估显示其领域内鲁棒性强,但跨领域泛化有限,提示未来需扩充微调数据以实现选择性干预而非机械套用。

原文摘要 · Abstract (English)

Large language models (LLMs) routinely face requests that should be refused, creating a trade-off between helpfulness and harm prevention. However, refusals themselves can be helpful. In high-risk interactions involving crisis, coercion, or escalating intent, blunt non-compliance may prevent direct harm while still failing to support the needs of the person behind the request. We present PsychoSafe, a psychologically-informed refusal framework that reframes refusal as structured supportive communication grounded in evidence-based intervention strategies. To develop PsychoSafe, we construct a corpus of 8019 prompt-response pairs spanning five psychologically salient risk domains and apply prompting and parameter-efficient fine-tuning to Qwen 3.5 27B. On a balanced validation set of 500 prompts, evaluated with an LLM judge and validated through human ratings, PsychoSafe prompting improves overall refusal quality by 28.1% over a generic baseline, with particularly strong gains in external resource referral (+46.8%) and psychological grounding (+34.8%), while preserving downstream performance on non-refusal tasks. Fine-tuning achieves near-perfect refusal and resource-referral rates but reduces response relevance. Additional evaluations on SORRY-Bench and XSTest show strong in-domain robustness but limited out-of-domain generalization, suggesting that future work should diversify fine-tuning data to help models apply interventions selectively rather than schematically.

大模型安全心理干预拒绝策略对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。