arXiv:2410.11195cs.CLcs.AI2024-10被引 5

用检索增强框架提升大模型法律判决预测能力

Athena: Retrieval-augmented Legal Judgment Prediction with Large Language Models

  • 构建罪名知识库,通过向量检索增强大模型输入
  • 在CAIL2018数据集上达顶尖性能,最高准确率95%
  • 揭示大模型'中间遗忘'现象,适合法律AI研究者

大型语言模型(LLMs)如ChatGPT、LLaMA和Claude已在众多领域广泛应用,包括法律场景。随着技术进步,提示工程(PE)作为连接LLMs与真实应用的接口日益受到关注。已有多种方法应对现实挑战,如少样本提示、思维链和检索增强生成(RAG)。然而,法律判决预测(LJP)中的RAG仍缺乏探索。为此,我们提出'Athena'框架,将RAG作为核心预处理组件,以提升LLMs在专业任务上的表现。Athena构建了包含罪名的知识库,并通过向量化实现语义检索。实验表明,Athena整体性能显著提升,在CAIL2018数据集上达到领先水平。对上下文窗口大小参数的消融研究进一步验证了LLMs的'中间遗忘'现象,且经适度超参数调优后,可达到最高95%的准确率。我们还分析了查询重写与数据分布的影响,为未来研究提供了可能方向。

原文摘要 · Abstract (English)

Recently, large language models (LLMs) like ChatGPT, LLaMA, and Claude have prevailed in countless domains, including legal scenarios. With LLMs' rapid technological progress, the development of prompt engineering (PE) as an interface between the LLMs and real-world applications has drawn the attention of all developers. Various PE methods have been proposed to overcome real-world challenges, such as few-shot prompting, chain-of-thought, and retrieval-augmented generation (RAG). However, RAG for legal judgment prediction (LJP) is still underexplored. To address this, we propose "Athena", a novel framework cultivating RAG as a core preprocess component to enhance LLMs' performance on specialized tasks. Athena constructs a knowledge base for accusations, attached with a semantic retrieval mechanism through vectorization. Our experiments show that Athena's overall performance has improved significantly, achieving state-of-the-art results on the CAIL2018 dataset. Our ablation study on the in-context window size parameter further reproduces LLMs' "lost-in-the-middle" phenomenon with a relative positional variation. And with moderate hyper-parameter-tuning, we can achieve at most 95% of accuracy accordingly. We also study the impact of query rewriting and data distribution, providing possible directions for future research based on former analyses.

法律AI检索增强大模型判决预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。