用大模型推理能力挖掘法律查询中的隐含概念,提升检索准确率
Exploiting LLMs' Reasoning Capability to Infer Implicit Concepts in Legal Information Retrieval
- 利用大模型逻辑推理能力补全法律查询中的隐含事实与术语
- 在COLIEE 2022和2023数据集上超越所有参赛团队最佳成绩
- 适合需要处理现实场景法律问题的研究者和开发者
法规检索是法律语言处理中的典型问题,具有广泛的实际应用价值。基于深度学习的检索方法虽已取得显著进展,但依赖语义与词汇关联的系统在面对涉及现实情境或非法律领域术语的查询时仍存在局限。本文通过利用大语言模型(LLMs)的逻辑推理能力,识别查询中提及情境相关的法律术语与事实,增强检索信息。所提系统结合术语扩展与查询重写策略,提升检索精度。在COLIEE 2022和COLIEE 2023数据集上的实验表明,大模型提供的额外知识可有效提升词汇与语义排序模型的表现。最终集成系统在两项竞赛中均优于所有参赛团队的最高水平。
原文摘要 · Abstract (English)
Statutory law retrieval is a typical problem in legal language processing, that has various practical applications in law engineering. Modern deep learning-based retrieval methods have achieved significant results for this problem. However, retrieval systems relying on semantic and lexical correlations often exhibit limitations, particularly when handling queries that involve real-life scenarios, or use the vocabulary that is not specific to the legal domain. In this work, we focus on overcoming this weaknesses by utilizing the logical reasoning capabilities of large language models (LLMs) to identify relevant legal terms and facts related to the situation mentioned in the query. The proposed retrieval system integrates additional information from the term--based expansion and query reformulation to improve the retrieval accuracy. The experiments on COLIEE 2022 and COLIEE 2023 datasets show that extra knowledge from LLMs helps to improve the retrieval result of both lexical and semantic ranking models. The final ensemble retrieval system outperformed the highest results among all participating teams in the COLIEE 2022 and 2023 competitions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。