小模型在法律场景下易因角色设定过度拒绝,可能引入偏见。
LLMs Prompted for Legal Context Object More: Overrefusal from Small On-Premises LLMs in Criminal Legal Context

- 测试小型本地部署模型在法律提示下的拒绝行为,发现角色设定显著提升拒绝率。
- 权威前缀使拒绝率提升2至20倍,远超无前缀基准。
- 适合关注AI法律应用伦理与偏见风险的研究者或从业者参考。
尽管大语言模型在法律领域的应用仍存在伦理与法律争议,法律从业者已开始尝试使用个人部署的小型LLM进行翻译与改写等任务。然而,即使这类看似无害的用途,也可能因模型对特定话题的选择性拒绝对案件处理速度引入偏差。为更好预判此类偏差,我们研究了几种最可能被用作本地设备助手的现代小型LLM,评估其在法律提示下的过量拒绝现象。结果显示,权威风格前缀(如“你作为国家最高法院助理”、“……辩护律师”)使拒绝率较无前缀基线系统性提升2–20倍;而一种已知的角色扮演越狱前缀则表现不一,在部分模型中大幅增加拒绝率,而在其他模型中几乎无影响。这一发现表明,小型本地可部署的LLM在真实机构用户可能自然引入的情境框架下表现不稳定,亟需进一步研究以减少偏见风险。
原文摘要 · Abstract (English)
While the validity of LLMs' use in the legal context remains subject to ethical and legal debate, legal professionals are already experimenting with personal LLMs, if only for translation and reformulation. However, even such a seemingly innocuous use can introduce biases through case processing speed if LLM assistants selectively refuse assistance on certain topics. To better anticipate such biases, we investigate several modern small LLMs that are most likely to be used as on-device assistants, to assess the impact of overrefusal on legal prompts. Surprisingly, we find that authority-style prefixes (``you are acting as an assistant of the national supreme court'', ``[...] defense lawyer'') systematically increase refusal rates by 2--20x over the no-prefix baseline, while a known role-play jailbreak prefix shows mixed effects, sharply increasing refusals in some models and barely shifting them in others. The finding suggests that small on-prem deployable LLMs are unstable under contextual framings that a real institutional user might naturally introduce, and further investigation is essential to minimize opportunities for bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。