首个专测台湾法律理解的LLM评估基准,填补本地化法律AI评测空白。
TW-LegalBench: Measuring Taiwanese Legal Understanding

- 构建覆盖5年、18个领域的超1.6万道选择题与14000+判例预测数据
- 顶尖模型通过律师资格考但不及法官/检察官水平,判罚引用法条能力弱
- 首次用评分细则量化评估法律作文,适合法律AI与司法科技研究者
大语言模型在多任务中表现优异,但在特定法域的法律推理能力仍待深入评估。本文提出TW-LegalBench,利用台湾公开的官方法律语料库,填补英美法系以英文资料和华语民法系以简体中文资料为主的评测空白。该基准包含三类任务:(1)涵盖五年、十八个专业领域的16,000余道多项选择题;(2)117道法律专业人士考试的开放性论述题,附官方评分细则;(3)超过14,000个涵盖数百罪名的法律判决预测实例。我们采用准确率评估选择题,基于评分细则点分解的LLM-as-Judge框架评估论述题,并以量刑准确率与法条引用率评估判决预测。结果表明,顶尖模型虽超过律师资格考试及格线(通过率11%),但未达法官与检察官标准(及格率1~2%)。在判决预测中,模型对罪名类型判断与刑期预测尚可,但难以准确引用具体法律条文。研究揭示:尽管大模型在资格考试中接近人类水平,可靠法律文本生成仍是挑战。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown impressive capabilities across diverse tasks, yet their performance on jurisdiction-specific legal reasoning remains underexplored. We present TW-LegalBench that utilizes Taiwanese legal system's rich official corpus open to the public to fill the gap in evaluating LLMs on Taiwanese law, among common-law benchmarks that focus on English sources and civil-law benchmarks focusing on sources of Simplified Chinese. TW-LegalBench comprises three task types: (1) over 16,000 multiple-choice questions (MCQs) across five years of official examinations in 18 professional domains; (2) 117 open-ended essay questions (OEQs) from examinations for legal professionals with official scoring rubrics; and (3) more than 14,000 legal judgment prediction (LJP) instances covering hundreds of crime categories. We evaluate 13 LLMs using accuracy for MCQs, a decomposed LLM-as-Judge framework based on the scoring rubric points for OEQs, and metrics for sentencing accuracy and statute citation for LJP. Our results reveal that top-performing models exceed the passing threshold for qualified lawyers (passing rate: 11%) but fall short of that for judges and prosecutors (passing rate: 1~2%). For LJP, while models demonstrate reasonable verdict type accuracy and sentence prediction capability, they struggle to cite exact legal articles. These findings highlight that reliable legal text generation remains challenging for LLMs, even though their performance on qualification examinations approaches human level.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。