arXiv:2512.11297cs.CL2025-12

构建日本企业法律任务开源基准,评估大模型长文本结构化输出能力

LegalRikai: Open Benchmark -- Benchmark for Complex Japanese Corporate Legal Tasks

  • 由律师主导设计4类真实企业法律任务,含100个需长文本输出的样本
  • 人类评估发现模型在文档级编辑上表现弱,常规短文本测试无法识别此缺陷
  • 自动化评估可有效筛选结果,适合专家资源有限的研究场景

本文介绍LegalRikai:一个新提出的开源基准,包含四个模拟日本企业法律实践的复杂任务。该基准由法律专业人士在律师监督下创建,共包含100个样本,要求生成长篇、结构化的输出,并针对多项实际标准进行评估。我们使用GPT-5、Gemini 2.5 Pro和Claude Opus 4.1等主流大模型进行了人工与自动评估。人类评估显示,抽象指令会引发不必要的修改,暴露出模型在文档级编辑上的不足,而这类问题在传统短文本任务中难以察觉。分析还表明,自动化评估在具有明确语言基础的标准上与人工判断高度一致,但结构一致性评估仍具挑战性。结果证明,自动化评估可在专家资源有限时作为有效筛选工具。本文提出一套数据集评估框架,以推动法律领域更贴近实践的研究。

原文摘要 · Abstract (English)

This paper introduces LegalRikai: Open Benchmark, a new benchmark comprising four complex tasks that emulate Japanese corporate legal practices. The benchmark was created by legal professionals under the supervision of an attorney. This benchmark has 100 samples that require long-form, structured outputs, and we evaluated them against multiple practical criteria. We conducted both human and automated evaluations using leading LLMs, including GPT-5, Gemini 2.5 Pro, and Claude Opus 4.1. Our human evaluation revealed that abstract instructions prompted unnecessary modifications, highlighting model weaknesses in document-level editing that were missed by conventional short-text tasks. Furthermore, our analysis reveals that automated evaluation aligns well with human judgment on criteria with clear linguistic grounding, and assessing structural consistency remains a challenge. The result demonstrates the utility of automated evaluation as a screening tool when expert availability is limited. We propose a dataset evaluation framework to promote more practice-oriented research in the legal domain.

法律AI文本生成评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。