arXiv:2604.13583cs.CLcs.AI2026-04中稿 · ICAIL 2026被引 1

打造一站式德国法律大模型评测平台,支持全流程协作与多维度评估。

BenGER Platform: A Collaborative Web Platform for End-to-End Benchmarking of German Legal Tasks

  • 构建集成任务设计、标注、模型运行与评估的Web平台
  • 支持词法、语义、事实及法官评分等多维指标评测
  • 适合法律专家与技术团队协作,提升评测透明度与可复现性

评估大语言模型在法律推理中的表现,需涵盖任务设计、专家标注、模型执行和基于指标的评估等多个环节。然而现实中这些步骤分散于不同平台与脚本中,限制了透明度、可复现性,也阻碍了非技术型法律专家参与。我们提出BenGER(德国法律基准)框架,一个开源的Web平台,整合了任务创建、协作标注、可配置的LLM运行以及基于词法、语义、事实和法官评分的评估功能。BenGER支持多机构项目,具备租户隔离与基于角色的访问控制,并可选地为标注者提供形成性、基于参考答案的反馈。我们将演示一个实时部署,展示端到端基准创建与分析流程。

原文摘要 · Abstract (English)

Evaluating large language models (LLMs) for legal reasoning requires workflows that span task design, expert annotation, model execution, and metric-based evaluation. In practice, these steps are split across platforms and scripts, limiting transparency, reproducibility, and participation by non-technical legal experts. We present the BenGER (Benchmark for German Law) framework, an open-source web platform that integrates task creation, collaborative annotation, configurable LLM runs, and evaluation with lexical, semantic, factual, and judge-based metrics. BenGER supports multi-organization projects with tenant isolation and role-based access control, and can optionally provide formative, reference-grounded feedback to annotators. We will demonstrate a live deployment showing end-to-end benchmark creation and analysis.

法律AI评测平台大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。