arXiv:2603.23646cs.CLcs.AI2026-03被引 1

首个评估前沿模型在瑞士监管任务表现的多语言基准,揭示大模型在合规场景中的真实短板。

Swiss-Bench SBP-002: A Frontier Model Comparison on Swiss Legal and Regulatory Tasks

  • 构建涵盖395个专家题目的三语多领域监管任务基准,覆盖7类任务与3种语言。
  • 顶级模型仅达38.2%正确率,监管问答等任务正确率低于9%,普遍存在幻觉问题。
  • 开源模型表现不俗,部分超越闭源模型,为监管场景模型选型提供实证参考。

尽管已有研究对大型语言模型在瑞士法律翻译(Niklaus等,2025)和学术法律推理(Fan等,2025)上进行了评测,但尚无基准评估前沿模型在实际瑞士监管合规任务中的表现。本文引入Swiss-Bench SBP-002,一个包含395个专家设计题目的三语基准,覆盖FINMA、Legal-CH、EFK三个瑞士监管领域,七类任务类型及德语、法语、意大利语三种语言。使用结构化三维评分框架,通过盲评三法官大模型小组(GPT-4o、Claude Sonnet 4、Qwen3-235B)进行多数投票聚合,加权卡帕系数为0.605;参考答案由独立人类法律专家在100个样本子集上验证,正确率73%,错误率为0%,实现完美法律准确性。结果揭示三个性能集群:A级(35-38%正确)、B级(26-29%)、C级(13-21%)。该基准难度极高,即使排名第一的模型(Qwen 3.5 Plus)也仅达38.2%正确率,47.3%错误,14.4%部分正确。任务类型差异显著:法律翻译与案例分析正确率达69-72%,而监管问答、幻觉检测与缺口分析均低于9%。在十款模型中(七款开源、三款闭源),一款开源模型领先,多款开源模型表现媲美或优于闭源模型。研究为零检索条件下评估前沿模型在瑞士监管任务中的能力提供了首个实证基准。

原文摘要 · Abstract (English)

While recent work has benchmarked large language models on Swiss legal translation (Niklaus et al., 2025) and academic legal reasoning from university exams (Fan et al., 2025), no existing benchmark evaluates frontier model performance on applied Swiss regulatory compliance tasks. I introduce Swiss-Bench SBP-002, a trilingual benchmark of 395 expert-crafted items spanning three Swiss regulatory domains (FINMA, Legal-CH, EFK), seven task types, and three languages (German, French, Italian), and evaluate ten frontier models from March 2026 using a structured three-dimension scoring framework assessed via a blind three-judge LLM panel (GPT-4o, Claude Sonnet 4, Qwen3-235B) with majority-vote aggregation and weighted kappa = 0.605, with reference answers validated by an independent human legal expert on a 100-item subset (73% rated Correct, 0% Incorrect, perfect Legal Accuracy). Results reveal three descriptive performance clusters: Tier A (35-38% correct), Tier B (26-29%), and Tier C (13-21%). The benchmark proves difficult: even the top-ranked model (Qwen 3.5 Plus) achieves only 38.2% correct, with 47.3% incorrect and 14.4% partially correct. Task type difficulty varies widely: legal translation and case analysis yield 69-72% correct rates, while regulatory Q&A, hallucination detection, and gap analysis remain below 9%. Within this roster (seven open-weight, three closed-source), an open-weight model leads the ranking, and several open-weight models match or outperform their closed-source counterparts. These findings provide an initial empirical reference point for assessing frontier model capability on Swiss regulatory tasks under zero-retrieval conditions.

法律AI监管合规多语言模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。