融合检索与大模型的法律推理框架,获法律文本蕴含任务冠军
NOWJ@COLIEE 2025: A Multi-stage Framework Integrating Embedding Models and Large Language Models for Legal Retrieval and Entailment
- 分阶段整合传统检索与大模型,先过滤后精排
- 法律蕴含任务取得F1 0.3195,排名第一
- 适合法律AI研究者和司法智能化开发者
本文介绍了NOWJ团队在COLIEE 2025竞赛全部五项任务中的方法与结果,重点突出法律案例蕴含任务(任务2)的进展。我们的综合方法系统性地结合了预排序模型(BM25、BERT、monoT5)、基于嵌入的语义表示(BGE-m3、LLM2Vec)以及先进的大语言模型(Qwen-2、QwQ-32B、DeepSeek-V3),用于摘要生成、相关性评分和上下文重排序。在任务2中,采用两阶段检索系统,结合词法-语义过滤与上下文化的大模型分析,以F1分数0.3195获得第一名。此外,在法律案例检索、法规检索、法律文本蕴含和法律判决预测等任务中,通过精心设计的集成策略与有效的提示推理方法,均展现出稳健性能。研究结果表明,传统信息检索技术与现代生成模型的混合架构具有巨大潜力,为未来法律信息处理提供了重要参考。
原文摘要 · Abstract (English)
This paper presents the methodologies and results of the NOWJ team's participation across all five tasks at the COLIEE 2025 competition, emphasizing advancements in the Legal Case Entailment task (Task 2). Our comprehensive approach systematically integrates pre-ranking models (BM25, BERT, monoT5), embedding-based semantic representations (BGE-m3, LLM2Vec), and advanced Large Language Models (Qwen-2, QwQ-32B, DeepSeek-V3) for summarization, relevance scoring, and contextual re-ranking. Specifically, in Task 2, our two-stage retrieval system combined lexical-semantic filtering with contextualized LLM analysis, achieving first place with an F1 score of 0.3195. Additionally, in other tasks--including Legal Case Retrieval, Statute Law Retrieval, Legal Textual Entailment, and Legal Judgment Prediction--we demonstrated robust performance through carefully engineered ensembles and effective prompt-based reasoning strategies. Our findings highlight the potential of hybrid models integrating traditional IR techniques with contemporary generative models, providing a valuable reference for future advancements in legal information processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。