arXiv:2412.14556cs.CL2024-12ACL被引 27

首个法律领域引用评估基准,提升大模型法律回答的准确性与可信度。

CitaLaw: Enhancing LLM with Citations in Legal Domain

  • 构建法律问答与判例参考库,支持大模型精准引用法条和判例。
  • 引入三段论评估方法,验证引用与回答的法律一致性,人工评分吻合度高。
  • 适用于法律AI研发者、司法科技从业者及法学教育研究者。

本文提出CitaLaw,首个用于评估大语言模型在法律领域生成合规回应并恰当引用的基准。该框架包含面向普通民众与专业人士的多样化法律问题,配套涵盖法律条文与判例的全面参考语料库。系统可从语料库中检索支持性引文,并将其与生成回答中的具体语句对齐。此外,我们设计了受三段论启发的评估方法,衡量引用与模型回答之间的法律一致性及其与用户问题的匹配程度。在2个通用领域和7个法律专用大模型上的实验表明,引入法律参考显著提升了回答质量。所提三段论评估方法与人工判断具有高度一致性。

原文摘要 · Abstract (English)

In this paper, we propose CitaLaw, the first benchmark designed to evaluate LLMs' ability to produce legally sound responses with appropriate citations. CitaLaw features a diverse set of legal questions for both laypersons and practitioners, paired with a comprehensive corpus of law articles and precedent cases as a reference pool. This framework enables LLM-based systems to retrieve supporting citations from the reference corpus and align these citations with the corresponding sentences in their responses. Moreover, we introduce syllogism-inspired evaluation methods to assess the legal alignment between retrieved references and LLM-generated responses, as well as their consistency with user questions. Extensive experiments on 2 open-domain and 7 legal-specific LLMs demonstrate that integrating legal references substantially enhances response quality. Furthermore, our proposed syllogism-based evaluation method exhibits strong agreement with human judgments.

法律AI大模型评估引用生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。