arXiv:2606.23375cs.CLcs.AI2026-06

测试并缓解大模型在多语种刑事法庭中的过度合规问题

Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts

论文配图:Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts
图 1 · 摘自论文原文
  • 构建多语言刑事法律文本评测集TF-RefusalBench,含5200个易触发拒绝的指令
  • 发现模型拒绝对策受模型类型、语言和文本内容共同影响,且免责声明会降低任务忠实度
  • 证明去除拒绝指令(abliteration)可几乎消除拒绝,同时保持任务性能

尽管大型语言模型在法律领域的广泛应用仍因可靠性与错误后果备受争议,但一些风险可控的特定应用已出现。瑞士联邦最高法院使用本地部署的小模型进行四种官方语言间的初步翻译与短文摘要。然而,在刑事司法场景中,此类应用面临挑战:案件材料常包含暴力与性犯罪的详细描述,导致模型安全机制被触发,产生拒绝响应与免责声明,干扰合法工作。为量化该现象,我们提出TF-RefusalBench,一个基于公开瑞士联邦最高法院裁决的多语言刑事法律翻译与摘要评测集,涵盖法语、德语、意大利语和英语共5200个提示。实验表明,过度对齐是多重因素作用的结果,受模型、提示及文本语言影响,其影响不能仅从拒绝率评估,因免责声明会显著降低输出忠实度。我们进一步评估了本地部署模型在刑事任务中的可行性,结果表明,虽然提示工程有效,但去除拒绝指令(abliteration)可在几乎不损害任务表现的前提下彻底消除拒绝。

原文摘要 · Abstract (English)

While the wider applicability of LLMs in the legal field is currently debated due to their reliability and the gravity of any errors, narrow uses with well-understood and mitigated risks have emerged. Notably the Swiss Federal Supreme Court uses small on-premises models for tentative translations and short-passage summarization across the four official languages. However, such usage is challenging in the context of Criminal Law. Since rulings and cases employees work on routinely can contain detailed descriptions of violent and sexual offenses, their legitimate work is compromised by refusals and disclaimers due to the activation of model guardrails (over-alignment). To measure this phenomenon, we introduce TF-RefusalBench, a multilingual benchmark for criminal-law translation and summarization derived from public Swiss Supreme Court rulings. TF-RefusalBench contains 5,200 total prompts across French, German, Italian, and English, corresponding to common task prompts and passages likely to trigger refusal. We then use TF-RefusalBench to show that over-alignment is a multifaceted phenomenon, influenced by the model and the prompt and text languages being processed, and that its impact cannot be evaluated solely from an over-refusal perspective, given the disclaimer's impact on task faithfulness. Finally, we evaluate approaches to enable on-premises LLMs for Criminal Law Tasks, demonstrating that while prompting can be effective, abliteration (refusal directions ablation) eliminates refusal with minimal impact on task performance.

大模型安全法律AI多语言过对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。