专精中国劳动法的AI模型与评测基准,提升法律AI准确性和实用性。
Chinese Labor Law Large Language Model Benchmark
- 构建专注中国劳动法的LLM模型及覆盖6类任务的评测集。
- 在多任务中超越GPT-4等通用与专用模型,平均性能领先12%以上。
- 方法可推广至其他法律领域,助力高可靠性法律AI落地。
大语言模型在特定领域应用中取得显著进展,尤其在法律领域。然而,通用模型如GPT-4在需要精确法律知识、复杂推理和上下文敏感性的细分领域表现不足。为此,我们提出LabourLawLLM——一款专用于中国劳动法的法律大模型,并构建LabourLawBench,一个涵盖法律条文引用、基于知识问答、案例分类、赔偿计算、命名实体识别和法律案例分析等六类任务的综合性评测基准。评估框架结合客观指标(如ROUGE-L、准确率、F1、soft-F1)与基于GPT-4的主观评分。实验表明,LabourLawLLM在各类任务中持续优于通用及现有法律专用模型。该方法不仅适用于劳动法,还可扩展至其他法律子领域,提升法律AI应用的准确性、可靠性与社会价值。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have led to substantial progress in domain-specific applications, particularly within the legal domain. However, general-purpose models such as GPT-4 often struggle with specialized subdomains that require precise legal knowledge, complex reasoning, and contextual sensitivity. To address these limitations, we present LabourLawLLM, a legal large language model tailored to Chinese labor law. We also introduce LabourLawBench, a comprehensive benchmark covering diverse labor-law tasks, including legal provision citation, knowledge-based question answering, case classification, compensation computation, named entity recognition, and legal case analysis. Our evaluation framework combines objective metrics (e.g., ROUGE-L, accuracy, F1, and soft-F1) with subjective assessment based on GPT-4 scoring. Experiments show that LabourLawLLM consistently outperforms general-purpose and existing legal-specific LLMs across task categories. Beyond labor law, our methodology provides a scalable approach for building specialized LLMs in other legal subfields, improving accuracy, reliability, and societal value of legal AI applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。