用关键段落替代全文,提升法律判例检索效率与准确性。
Segment First, Retrieve Better: Realistic Legal Search via Rhetorical Role-Based Queries
- 基于修辞角色提取重要段落,减少输入信息依赖。
- 在IL-PCR和COLIEE 2025上优于传统方法,提升检索精度。
- 适合仅掌握部分案情的法律实务人员使用。
法律判例检索是普通法体系的核心,遵循先例原则(stare decisis),要求判决一致性。然而,法律文件日益复杂且数量庞大,传统检索方法面临挑战。TraceRetriever 模拟真实法律搜索场景,仅依赖有限案件信息,通过提取具有修辞意义的关键段落而非完整文档进行检索。其流程整合了 BM25、向量数据库与交叉编码器模型,采用倒数排名融合(Reciprocal Rank Fusion)合并初筛结果,再进行最终重排序。修辞标注由在印度判例上训练的层级双向LSTM-CRF分类器生成。在 IL-PCR 与 COLIEE 2025 数据集上的评估表明,该方法有效应对文档体量增长问题,同时契合实际检索约束,为仅有部分案情知识的法律研究提供可靠、可扩展的判例检索基础。
原文摘要 · Abstract (English)
Legal precedent retrieval is a cornerstone of the common law system, governed by the principle of stare decisis, which demands consistency in judicial decisions. However, the growing complexity and volume of legal documents challenge traditional retrieval methods. TraceRetriever mirrors real-world legal search by operating with limited case information, extracting only rhetorically significant segments instead of requiring complete documents. Our pipeline integrates BM25, Vector Database, and Cross-Encoder models, combining initial results through Reciprocal Rank Fusion before final re-ranking. Rhetorical annotations are generated using a Hierarchical BiLSTM CRF classifier trained on Indian judgments. Evaluated on IL-PCR and COLIEE 2025 datasets, TraceRetriever addresses growing document volume challenges while aligning with practical search constraints, reliable and scalable foundation for precedent retrieval enhancing legal research when only partial case knowledge is available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。