arXiv:2606.18158cs.CYcs.AI2026-06

填补法律AI评估空白,首次测试大模型的判例推理能力

The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act

  • 构建首个针对欧盟法判例推理的基准测试
  • 发现现有大模型在核心法律推理任务上表现不足
  • 为欧盟高风险AI合规提供可操作评估标准

大型语言模型如今已能生成至少达到中位质量的法律文本,但现有评估基准无法衡量它们是否具备判例推理能力——这正是法律工作的解释性核心,而非当前多数法律AI评估所关注的辅助性、事务性任务。这一评估空白不仅是方法论问题,更是法律问题:欧盟《人工智能法案》要求司法领域使用的高风险AI必须满足‘适当准确性’,但若无具备判例推理能力的评估基准,该要求将无法获得实际操作意义。

原文摘要 · Abstract (English)

Large language models now produce legal text of at least median quality, yet no existing benchmark can evaluate whether they perform doctrinal legal reasoning, which forms the interpretive core of legal work, rather than the ancillary, paralegal tasks that most current legal-AI evaluations measure. This measurement gap is not only methodological but legal: the EU AI Act makes "appropriate accuracy" a binding requirement for high-risk AI used in the judicial domain, yet that requirement cannot acquire operational content without the very doctrinal-reasoning benchmark the field lacks.

法律AI判例推理欧盟法评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。