研究大模型拒答对用户满意度的影响,发现伦理拒答最让人不满。
LLM Content Moderation and User Satisfaction: Evidence from Response Refusals in Chatbot Arena
- 用细调的RoBERTa模型区分伦理与技术性拒答
- 伦理拒答导致胜率显著下降,用户更不满意
- 详细且贴合语境的拒答能缓解负面评价
大模型安全与伦理对齐备受关注,但内容审核对用户满意度的影响仍不清楚。尤其缺乏对模型因伦理原因拒绝回答时用户反应的研究。我们通过分析近5万次来自Chatbot Arena平台的模型对比数据,该平台让用户在成对响应中选择偏好模型,提供了真实世界用户偏好的大规模观察场景。采用基于RoBERTa的拒答分类器(在人工标注数据集上微调),我们区分了由伦理问题和非伦理因素导致的拒答。结果表明:伦理拒答带来的胜率损失显著高于技术拒答和正常回应,说明用户对因伦理原因拒答尤为不满。但这一惩罚并非恒定:当提示涉及高度敏感内容(如违法信息)或拒答措辞详尽且上下文一致时,用户评价更积极。这揭示了大模型设计中的核心矛盾——安全对齐行为可能违背用户期待,亟需更适应上下文和表达方式的动态审核策略。
原文摘要 · Abstract (English)
LLM safety and ethical alignment are widely discussed, but the impact of content moderation on user satisfaction remains underexplored. In particular, little is known about how users respond when models refuse to answer a prompt-one of the primary mechanisms used to enforce ethical boundaries in LLMs. We address this gap by analyzing nearly 50,000 model comparisons from Chatbot Arena, a platform where users indicate their preferred LLM response in pairwise matchups, providing a large-scale setting for studying real-world user preferences. Using a novel RoBERTa-based refusal classifier fine-tuned on a hand-labeled dataset, we distinguish between refusals due to ethical concerns and technical limitations. Our results reveal a substantial refusal penalty: ethical refusals yield significantly lower win rates than both technical refusals and standard responses, indicating that users are especially dissatisfied when models decline a task for ethical reasons. However, this penalty is not uniform. Refusals receive more favorable evaluations when the underlying prompt is highly sensitive (e.g., involving illegal content), and when the refusal is phrased in a detailed and contextually aligned manner. These findings underscore a core tension in LLM design: safety-aligned behaviors may conflict with user expectations, calling for more adaptive moderation strategies that account for context and presentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。