测试了'反向请求'对大模型的影响力,发现效果因模型而异。
Door-in-the-Face Requests and Refusal Behaviour in Large Language Models
- 先拒大请求再提小请求,观察模型是否更配合。
- Anthropic模型在拒后小请求成功率升至65.8%,其他模型反而下降15.5~23.0点。
- 能否成功取决于请求内容,解释类请求比指令类更易被接受。
人类中‘门内之面’技巧有效:先提出大请求被拒,再提小请求更容易成功。我们测试了来自三家厂商的九个上线模型,均先拒绝一个大请求,再提出相同主题的小请求,对比直接提问的响应率。结果因模型而异:Anthropic的Opus 5在拒后小请求的响应率达65.8%,高于直接提问的29.3%;而OpenAI和Google的前沿模型及Haiku 4.5则出现反效果,响应率下降15.5至23.0个百分点。控制实验显示,无关话题的拒绝无此效应,说明‘拒绝行为本身’起作用,但模型反应机制不同。该技巧无法迁移至公开基准中的拒绝样本。关键在于请求内容:将265个被拒的指令改写为同主题解释请求,其中263例成功解除拒绝。表明人类影响策略需逐模型适配。
原文摘要 · Abstract (English)
Does the door-in-the-face technique work on language models? In humans, a large request that is refused makes a smaller follow-up request more likely to be granted. We test this on nine production models from three providers: each model refuses a large request, then receives a smaller version of the same request, and we compare its compliance with asking directly. The answer depends on the model. On Anthropic's frontier models the technique works: Opus 5 answers the smaller request 65.8% of the time after refusing the larger one, against 29.3% when asked directly. On the frontier models of OpenAI and Google, and on Haiku 4.5, it backfires, lowering compliance by 15.5 to 23.0 points. A control locates the effect: a refused large request on an unrelated topic does less than the related one on all nine models, so the concession itself matters everywhere, while the reaction to having just refused something differs by model family. The technique does not transfer to refusals drawn from public benchmarks. What decides whether a retreat can work is what the request asks for: rewriting 265 refused requests for usable instructions into requests for explanations of the same topic removed the refusal in 263 cases. Human influence techniques port to language models one model family at a time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。