测试大模型在法律领域的事实准确性,发现合理容错和拒答能显著提升可靠性。
The Factuality of Large Language Models in the Legal Domain
- 设计法律案例与法规的多样化事实问答数据集,支持多维度评估。
- 使用模糊匹配等方法后,模型准确率从63%提升至81%。
- 适合法律AI研究者,尤其关注模型可靠性与可解释性的人群。
本文研究大语言模型(LLMs)作为法律领域知识库的事实性,在真实使用场景下允许答案存在合理差异,并允许模型在不确定时拒绝回答。首先,我们构建了一个涵盖判例法与立法的多样化事实问题数据集。随后,利用该数据集在不同评估方法(精确匹配、别名匹配、模糊匹配)下评估多个LLMs。结果表明,采用别名与模糊匹配方法后性能显著提升。进一步探索拒答机制与上下文示例的影响,发现两者均能提高精确度。最后,通过在法律文档上进行额外预训练(如SaulLM),事实精确率从63%提升至81%。
原文摘要 · Abstract (English)
This paper investigates the factuality of large language models (LLMs) as knowledge bases in the legal domain, in a realistic usage scenario: we allow for acceptable variations in the answer, and let the model abstain from answering when uncertain. First, we design a dataset of diverse factual questions about case law and legislation. We then use the dataset to evaluate several LLMs under different evaluation methods, including exact, alias, and fuzzy matching. Our results show that the performance improves significantly under the alias and fuzzy matching methods. Further, we explore the impact of abstaining and in-context examples, finding that both strategies enhance precision. Finally, we demonstrate that additional pre-training on legal documents, as seen with SaulLM, further improves factual precision from 63% to 81%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。