arXiv:2502.16357cs.CL2025-02被引 3

首个葡萄牙法律大模型评测基准,基于真实考题生成高质量测试集。

LegalBench.PT: A Benchmark for Portuguese Law

  • 用GPT-4o将真实法律考试题转为多选、判断等题型。
  • 经专业律师审核,模型在该基准上表现接近人类水平。
  • 适合评估法律大模型能力或研究语言偏差的研究者使用。

大型语言模型在法律领域的应用推动了多司法管辖区和语言的基准建设。然而,目前尚无专为葡萄牙法律体系设计的基准。本文提出LegalBench.PT,首个覆盖葡萄牙法律核心领域的综合性法律评测基准。我们首先收集真实法律考试中的长文本问答,再利用GPT-4o将其转换为多项选择、是非判断和配对题型。生成后,通过筛选与处理提升数据质量,并由法律专业人士对样例进行验证以确保准确性与相关性。尽管题目为合成生成,但其基于人类出题且经严格过滤,可作为评估大模型法律知识与推理能力的可靠基准。我们还评估了主流大模型在LegalBench.PT上的表现,并分析GPT-4o回答中的潜在偏差。同时,对葡萄牙律师在部分题目上的表现进行测试,建立模型对比基线并验证基准有效性。

原文摘要 · Abstract (English)

The recent application of LLMs to the legal field has spurred the creation of benchmarks across various jurisdictions and languages. However, no benchmark has yet been specifically designed for the Portuguese legal system. In this work, we present LegalBench.PT, the first comprehensive legal benchmark covering key areas of Portuguese law. To develop LegalBench.PT, we first collect long-form questions and answers from real law exams, and then use GPT-4o to convert them into multiple-choice, true/false, and matching formats. Once generated, the questions are filtered and processed to improve the quality of the dataset. To ensure accuracy and relevance, we validate our approach by having a legal professional review a sample of the generated questions. Although the questions are synthetically generated, we show that their basis in human-created exams and our rigorous filtering and processing methods applied result in a reliable benchmark for assessing LLMs' legal knowledge and reasoning abilities. Finally, we evaluate the performance of leading LLMs on LegalBench.PT and investigate potential biases in GPT-4o's responses. We also assess the performance of Portuguese lawyers on a sample of questions to establish a baseline for model comparison and validate the benchmark.

法律AI评测基准葡萄牙语大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。