区分法律文本生成中的漏洞与幻觉,提升评估准确性
Gaps or Hallucinations? Gazing into Machine-Generated Legal Analysis for Fine-grained Text Evaluations
- 用'漏洞'替代'幻觉'定义机器生成文本的差异,更中立合理
- 构建细粒度检测器,测试集上F1达67%,精度80%
- 发现主流LLM生成的法律分析约80%含各类幻觉
大型语言模型在法律分析写作辅助方面展现出潜力,但常产生难以被非专业人士识别的幻觉,现有文本评估指标也难以应对。本文提出:何时可认为机器生成的法律分析可接受?我们引入‘漏洞’这一中性概念,指代人工撰写与机器生成法律分析之间的差异,不等同于错误。基于法律专家协作,针对Hou等人(2024b)提出的CLERC生成任务,构建分类体系、细粒度漏洞检测器,并发布可用于自动评估的标注数据集。最佳检测器在测试集上达到67% F1分数和80%精确率。将该检测器作为自动化评估指标应用于当前顶尖大模型生成的法律分析,发现约80%存在不同类型的幻觉。
原文摘要 · Abstract (English)
Large Language Models (LLMs) show promise as a writing aid for professionals performing legal analyses. However, LLMs can often hallucinate in this setting, in ways difficult to recognize by non-professionals and existing text evaluation metrics. In this work, we pose the question: when can machine-generated legal analysis be evaluated as acceptable? We introduce the neutral notion of gaps, as opposed to hallucinations in a strict erroneous sense, to refer to the difference between human-written and machine-generated legal analysis. Gaps do not always equate to invalid generation. Working with legal experts, we consider the CLERC generation task proposed in Hou et al. (2024b), leading to a taxonomy, a fine-grained detector for predicting gap categories, and an annotated dataset for automatic evaluation. Our best detector achieves 67% F1 score and 80% precision on the test set. Employing this detector as an automated metric on legal analysis generated by SOTA LLMs, we find around 80% contain hallucinations of different kinds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。