arXiv:2505.17267cs.CL2025-05EMNLP被引 22

评测大模型在希腊法律题中的引文推理能力,发现顶尖模型仍不及顶级专家。

GreekBarBench: A Challenging Benchmark for Free-Text Legal Reasoning and Citations

  • 设计三维评分体系+大模型判卷,解决自由文本评估难题。
  • 13个模型最佳表现超平均专家,但未达95%顶尖专家水平。
  • 适合法律AI、评测方法研究者参考,推动司法大模型发展。

我们提出GreekBarBench,一个基于希腊律师考试的基准测试,涵盖五个法律领域,要求回答时引用法条和判例事实。为应对自由文本评估挑战,我们设计了三维评分系统,并结合大模型作为评判者。同时构建元评估基准,检验大模型判官与人类专家评价的相关性,发现简单的片段匹配评分标准能显著提升一致性。对13个专有及开源大模型的系统评估显示,尽管最优模型超过平均专家得分,但仍未达到专家群体第95百分位水平。

原文摘要 · Abstract (English)

We introduce GreekBarBench, a benchmark that evaluates LLMs on legal questions across five different legal areas from the Greek Bar exams, requiring citations to statutory articles and case facts. To tackle the challenges of free-text evaluation, we propose a three-dimensional scoring system combined with an LLM-as-a-judge approach. We also develop a meta-evaluation benchmark to assess the correlation between LLM-judges and human expert evaluations, revealing that simple, span-based rubrics improve their alignment. Our systematic evaluation of 13 proprietary and open-weight LLMs shows that even though the best models outperform average expert scores, they fall short of the 95th percentile of experts.

法律AI评测基准大模型判卷

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。