评测大模型在多语言法律推理中的表现,发现准确率普遍低于50%。
Evaluating the Limits of Large Language Models in Multilingual Legal Reasoning
- 用大模型做裁判,评估多语言法律任务性能
- 法律推理准确率常低于50%,远低于通用任务的70%
- 模型对提示和攻击敏感,英文非绝对优势
在大型语言模型(LLMs)主导的时代,理解其在高风险领域如法律中的能力与局限至关重要。尽管像Meta的LLaMA、OpenAI的ChatGPT、Google的Gemini、DeepSeek等模型正被越来越多地应用于法律工作流程,但它们在多语言、跨司法管辖区及对抗性环境下的表现仍缺乏充分研究。本文评估了LLaMA和Gemini在多语言法律与非法律基准上的表现,并通过字符级和词级扰动测试其对抗鲁棒性。采用大模型作为裁判的方法进行人类对齐评估。此外,提出一个开源、模块化的评估流水线,支持任意组合的模型与数据集在法律任务(包括分类、摘要、开放问答和通用推理)上的多语言、多任务基准测试。结果表明,法律任务对大模型构成显著挑战,在LEXam等法律推理基准上准确率常低于50%,而通用任务如XNLI上则超过70%。尽管英语通常表现更稳定,但并不总带来更高准确率;提示敏感性和对抗脆弱性在各语言中均存在。同时发现,语言与英语的句法相似性与其性能呈正相关。还观察到LLaMA弱于Gemini,后者在相同任务上平均领先约24个百分点。尽管新模型有所进步,但在关键的多语言法律应用中仍难实现可靠部署。
原文摘要 · Abstract (English)
In an era dominated by Large Language Models (LLMs), understanding their capabilities and limitations, especially in high-stakes fields like law, is crucial. While LLMs such as Meta's LLaMA, OpenAI's ChatGPT, Google's Gemini, DeepSeek, and other emerging models are increasingly integrated into legal workflows, their performance in multilingual, jurisdictionally diverse, and adversarial contexts remains insufficiently explored. This work evaluates LLaMA and Gemini on multilingual legal and non-legal benchmarks, and assesses their adversarial robustness in legal tasks through character and word-level perturbations. We use an LLM-as-a-Judge approach for human-aligned evaluation. We moreover present an open-source, modular evaluation pipeline designed to support multilingual, task-diverse benchmarking of any combination of LLMs and datasets, with a particular focus on legal tasks, including classification, summarization, open questions, and general reasoning. Our findings confirm that legal tasks pose significant challenges for LLMs with accuracies often below 50% on legal reasoning benchmarks such as LEXam, compared to over 70% on general-purpose tasks like XNLI. In addition, while English generally yields more stable results, it does not always lead to higher accuracy. Prompt sensitivity and adversarial vulnerability is also shown to persist across languages. Finally, a correlation is found between the performance of a language and its syntactic similarity to English. We also observe that LLaMA is weaker than Gemini, with the latter showing an average advantage of about 24 percentage points across the same task. Despite improvements in newer LLMs, challenges remain in deploying them reliably for critical, multilingual legal applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。