构建瑞士多语言法律翻译基准,评估大模型在真实法律文本上的表现。
SwiLTra-Bench: The Swiss Legal Translation Benchmark
- 构建18万+对齐的瑞士多语法律文本数据集,覆盖四国官方语言及英语。
- 顶尖大模型零样本表现优于微调的开源模型,尤其在法律条文上更准确。
- 提出专业评估系统SwiLTra-Judge,与人类专家判断高度一致。
瑞士因四种官方语言和多语言法律文件要求,法律翻译至关重要。传统上依赖兼具法律知识与翻译能力的专业人士,造成瓶颈并影响司法可及性。为此,我们推出SwiLTra-Bench,一个包含超过18万对齐瑞士法律翻译对的综合性多语言基准,涵盖法律条文、导言和新闻稿,覆盖所有瑞士官方语言及英语,用于评估基于大模型的翻译系统。系统性评估显示,前沿模型在各类文档中表现更优,而专用翻译系统在法律条文上表现突出,但在导言部分表现较差。通过严格测试与人工专家验证,我们证明微调开源小模型虽能显著提升质量,但仍逊于最佳零样本提示的前沿模型(如Claude-3.5-Sonnet)。此外,我们提出了与人类专家评估高度一致的专用评估系统SwiLTra-Judge。
原文摘要 · Abstract (English)
In Switzerland legal translation is uniquely important due to the country's four official languages and requirements for multilingual legal documentation. However, this process traditionally relies on professionals who must be both legal experts and skilled translators -- creating bottlenecks and impacting effective access to justice. To address this challenge, we introduce SwiLTra-Bench, a comprehensive multilingual benchmark of over 180K aligned Swiss legal translation pairs comprising laws, headnotes, and press releases across all Swiss languages along with English, designed to evaluate LLM-based translation systems. Our systematic evaluation reveals that frontier models achieve superior translation performance across all document types, while specialized translation systems excel specifically in laws but under-perform in headnotes. Through rigorous testing and human expert validation, we demonstrate that while fine-tuning open SLMs significantly improves their translation quality, they still lag behind the best zero-shot prompted frontier models such as Claude-3.5-Sonnet. Additionally, we present SwiLTra-Judge, a specialized LLM evaluation system that aligns best with human expert assessments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。