用自动化方法评估大模型生成法律论点时的可靠性与拒答能力。
Measuring Faithfulness and Abstention: An Automated Pipeline for Evaluating LLM-Generated 3-ply Case-Based Legal Arguments
- 通过外部大模型提取生成论点中的要素,对比原始案例数据验证真实性。
- 90%以上模型在可生成论点任务中不编造内容,但常遗漏关键事实要素。
- 多数模型无法在无共同事实时拒绝生成,反而胡编乱造,不适用于法律场景。
大语言模型在复杂法律任务如论点生成方面展现出潜力,但其可靠性仍存疑。本文基于前期人工评估工作,提出一种自动化评估管道,专门评测大模型生成三段式法律论点的表现,重点关注真实性(无幻觉)、要素利用和合理拒答能力。将幻觉定义为生成输入案例中不存在的要素,拒答指模型在无事实依据时按指令停止生成。该方法使用外部大模型从生成论点中提取要素,并与输入案例三元组(当前案与两个判例)中的真实要素比对。在三个难度递增的任务上测试了八种大模型:1)生成标准三段式论点;2)交换判例角色生成;3)识别因缺乏共通要素而无法生成论点,主动拒答。结果显示,当前模型在可行任务(测试1&2)中幻觉率低于10%,但普遍未能充分使用所有相关要素;在拒答测试(测试3)中,多数模型未遵循指令,仍生成虚假论点。该自动化管道为评估关键行为提供了可扩展方案,凸显提升要素利用与鲁棒拒答能力的必要性,以保障法律场景下的可靠部署。
原文摘要 · Abstract (English)
Large Language Models (LLMs) demonstrate potential in complex legal tasks like argument generation, yet their reliability remains a concern. Building upon pilot work assessing LLM generation of 3-ply legal arguments using human evaluation, this paper introduces an automated pipeline to evaluate LLM performance on this task, specifically focusing on faithfulness (absence of hallucination), factor utilization, and appropriate abstention. We define hallucination as the generation of factors not present in the input case materials and abstention as the model's ability to refrain from generating arguments when instructed and no factual basis exists. Our automated method employs an external LLM to extract factors from generated arguments and compares them against the ground-truth factors provided in the input case triples (current case and two precedent cases). We evaluated eight distinct LLMs on three tests of increasing difficulty: 1) generating a standard 3-ply argument, 2) generating an argument with swapped precedent roles, and 3) recognizing the impossibility of argument generation due to lack of shared factors and abstaining. Our findings indicate that while current LLMs achieve high accuracy (over 90%) in avoiding hallucination on viable argument generation tests (Tests 1 & 2), they often fail to utilize the full set of relevant factors present in the cases. Critically, on the abstention test (Test 3), most models failed to follow instructions to stop, instead generating spurious arguments despite the lack of common factors. This automated pipeline provides a scalable method for assessing these crucial LLM behaviors, highlighting the need for improvements in factor utilization and robust abstention capabilities before reliable deployment in legal settings. Link: https://lizhang-aiandlaw.github.io/An-Automated-Pipeline-for-Evaluating-LLM-Generated-3-ply-Case-Based-Legal-Arguments/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。