用新指标评估法律大模型,发现高级检索增强系统可将幻觉降至0.2%以下。
Reliability by design: quantifying and eliminating fabrication risk in LLMs. From generative to consultative AI: a comparative analysis in the legal domain and lessons for high-stakes knowledge bases
- 提出双指标量化幻觉,对比三类AI在法律任务中的表现
- 基础检索增强使错误率显著下降,高级版本近乎消除虚构事实
- 适合法律、医疗等高风险领域从业者参考,强调可追溯与验证
本文研究如何让大语言模型在高风险法律工作中保持可靠,以减少幻觉问题。区分三种AI范式:(1)独立生成模型(“创意预言家”),(2)基础检索增强系统(“专家档案员”),(3)端到端优化的高级RAG系统(“严谨档案员”)。作者引入两个可靠性指标——假引用率(FCR)和虚构事实率(FFR),对12个LLM在75项法律任务中生成的2700条司法风格回答进行专家双盲评审。结果显示,独立模型不适用于专业场景(FCR超30%),基础RAG虽显著降低错误但仍存在明显误引;而采用嵌入微调、重排序和自纠错等技术的高级RAG,将虚构事实率降至不足0.2%。研究结论指出,可信法律AI需以验证和可追溯性为核心的检索架构,并提供可推广至其他高风险领域的评估框架。
原文摘要 · Abstract (English)
This paper examines how to make large language models reliable for high-stakes legal work by reducing hallucinations. It distinguishes three AI paradigms: (1) standalone generative models ("creative oracle"), (2) basic retrieval-augmented systems ("expert archivist"), and (3) an advanced, end-to-end optimized RAG system ("rigorous archivist"). The authors introduce two reliability metrics -False Citation Rate (FCR) and Fabricated Fact Rate (FFR)- and evaluate 2,700 judicial-style answers from 12 LLMs across 75 legal tasks using expert, double-blind review. Results show that standalone models are unsuitable for professional use (FCR above 30%), while basic RAG greatly reduces errors but still leaves notable misgrounding. Advanced RAG, using techniques such as embedding fine-tuning, re-ranking, and self-correction, reduces fabrication to negligible levels (below 0.2%). The study concludes that trustworthy legal AI requires rigor-focused, retrieval-based architectures emphasizing verification and traceability, and provides an evaluation framework applicable to other high-risk domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。