arXiv:2509.08075cs.CL2025-09被引 7

研究发现,角色提示会误导模型错误拒绝请求,但影响程度被高估。

No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models

  • 用蒙特卡洛方法高效量化15类社会身份对拒绝对话的影响
  • 越强大的模型受角色提示影响越小,但任务和模型选择更关键
  • 安全机制可能存在偏见,适合关注模型公平性的研究者阅读

大型语言模型(LLMs)正日益个性化,但可能引发意外副作用。近期研究表明,角色提示可能导致模型错误拒绝用户请求,但尚未有研究全面量化该问题。为此,我们评估了基于性别、种族、宗教和残障等15类社会人口学角色提示对错误拒绝率的影响。为控制其他变量,测试了16种不同模型、3项任务(自然语言推理、礼貌性、冒犯性分类)及9种提示改写方式。提出一种基于蒙特卡洛的样本高效量化方法。结果显示,随着模型能力提升,角色提示对拒绝率的影响逐渐减弱。某些社会身份在特定模型中会增加错误拒绝,暗示对齐策略或安全机制存在潜在偏见。然而,模型选择与任务类型对错误拒绝影响显著,尤其在敏感内容任务中。结果表明,角色提示的影响可能被夸大,实际原因或来自其他因素。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly integrated into our daily lives and personalized. However, LLM personalization might also increase unintended side effects. Recent work suggests that persona prompting can lead models to falsely refuse user requests. However, no work has fully quantified the extent of this issue. To address this gap, we measure the impact of 15 sociodemographic personas (based on gender, race, religion, and disability) on false refusal. To control for other factors, we also test 16 different models, 3 tasks (Natural Language Inference, politeness, and offensiveness classification), and nine prompt paraphrases. We propose a Monte Carlo-based method to quantify this issue in a sample-efficient manner. Our results show that as models become more capable, personas impact the refusal rate less and less. Certain sociodemographic personas increase false refusal in some models, which suggests underlying biases in the alignment strategies or safety mechanisms. However, we find that the model choice and task significantly influence false refusals, especially in sensitive content tasks. Our findings suggest that persona effects have been overestimated, and might be due to other factors.

大模型角色提示偏见检测拒绝行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。