让大模型推荐更安全,能自动避开用户的个人敏感点。
SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems
- 引入个性化安全约束,训练时动态识别并避开用户敏感话题。
- 在新数据集上将安全违规率降低96.5%,推荐效果仍优秀。
- 适合关注大模型伦理安全、对话推荐系统的研究者和开发者。
当前基于大语言模型的对话推荐系统主要优化推荐准确率和用户满意度。我们发现一个被忽视的漏洞:当系统从对话中隐式推断出用户的个性化安全敏感性(如创伤触发点、自残史或恐惧症)但未在推荐中加以尊重时,推荐结果可能对用户造成负面影响。为此,我们首次提出个性化对话推荐系统安全问题,并构建SafeRec基准数据集,用于系统评估大模型推荐在个体化约束下的安全风险。为解决该问题,我们提出SafeCRS框架,融合安全监督微调(Safe-SFT)与安全组奖励解耦归一化策略优化(Safe-GDPO),联合优化推荐质量与个性化安全对齐。在SafeRec上的大量实验表明,相比最强推荐质量基线,SafeCRS将安全违规率降低高达96.5%,同时保持良好的推荐性能。警告:本文包含可能有害或冒犯性内容。
原文摘要 · Abstract (English)
Current LLM-based conversational recommender systems (CRS) primarily optimize recommendation accuracy and user satisfaction. We identify an underexplored vulnerability in which recommendation outputs may negatively impact users by violating personalized safety constraints, when individualized safety sensitivities -- such as trauma triggers, self-harm history, or phobias -- are implicitly inferred from the conversation but not respected during recommendation. We formalize this challenge as personalized CRS safety and introduce SafeRec, a new benchmark dataset designed to systematically evaluate safety risks in LLM-based CRS under user-specific constraints. To further address this problem, we propose SafeCRS, a safety-aware training framework that integrates Safe Supervised Fine-Tuning (Safe-SFT) with Safe Group reward-Decoupled Normalization Policy Optimization (Safe-GDPO) to jointly optimize recommendation quality and personalized safety alignment. Extensive experiments on SafeRec demonstrate that SafeCRS reduces safety violation rates by up to 96.5% relative to the strongest recommendation-quality baseline while maintaining competitive recommendation quality. Warning: This paper contains potentially harmful and offensive content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。