为大模型拒绝行为建立基于语用学的评估体系,揭示其拒绝方式可能引发用户不适。
You Shouldn't Have Asked: A Pragmatics-Inspired Taxonomy for Evaluating LLM Refusals
- 基于语用学理论构建首个大模型拒绝行为分类框架
- 多数模型拒绝时态度强硬且带有道德评判,较少进行人际修复
- 适合关注模型交互安全与用户体验的研究者和开发者
拒绝在语用学中被视为一种威胁面子的行为,可能损害请求者的社会自我形象。大语言模型(LLMs)越来越倾向于拒绝不安全或不当请求,但若未能妥善管理这种互动成本,反而可能伤害用户。现有研究多将模型不配合视为安全对齐的结果,却缺乏评估模型在不同有害情境下是否恰当拒绝的方法。为此,我们提出了首个基于语用学理论的LLM拒绝行为分类体系。将其应用于16个现代大模型在14类危害情境下的回应,发现尽管各模型拒绝方式存在差异,但总体上表现为明确且强道德评价性,互动修复主要通过提供更安全替代方案实现,而非人际层面的面子维护。这一模式在敏感危害场景中尤为关键——过度负面表述可能使用户感到羞辱或激怒,背离安全拒绝的初衷。因此,我们呼吁在对齐评估中不仅关注模型是否拒绝有害请求,还应考察其拒绝方式是否具备情境适应性和社交责任意识。
原文摘要 · Abstract (English)
Refusals are often treated as face-threatening acts in pragmatics because they can challenge the requester's socially claimed self-image. Large language models (LLMs) are increasingly trained to refuse unsafe and inappropriate requests, and these refusals may harm users when models fail to manage this interactional cost properly. While existing work has mainly approached LLM non-compliance as a safety-alignment outcome, it does not provide a way to evaluate whether LLMs refuse appropriately across different harmful contexts. To study this question, we propose (to our knowledge) the first taxonomy of LLM refusals that is grounded in pragmatic theory. Applying this taxonomy to responses from 16 modern LLMs across 14 harm categories, we find that although models differ in how they refuse, their refusals are overall explicit and strongly morally evaluative, with interactional repair occurring mainly through offering or providing safer alternatives instead of interpersonal facework. This pattern is especially consequential in sensitive harm contexts, where overuse of negative framing may make users feel shamed or provoked, undermining the purpose of safe non-compliance. We therefore call for alignment evaluation that considers not only whether models refuse harmful requests, but also whether they refuse in ways that are contextually adaptive and socially accountable for the interactional consequences of saying no.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。