arXiv:2411.07739cs.AIcs.IR2024-11被引 8

用多层嵌入提升法律文本检索精度,支持细粒度到整章的精准问答。

Unlocking Legal Knowledge with Multi-Layered Embedding-Based Retrieval

  • 为法律条文、段落、条款等不同层级构建嵌入向量,实现多粒度表示。
  • 在巴西宪法与立法文本上验证,能准确响应具体条款或整体章节查询。
  • 方法可推广至其他法律体系及层级化文本领域,如政策文件、标准文档。

本文针对法律知识的复杂性,提出一种基于多层嵌入的检索方法,用于法律与立法文本。不仅为单个条文,还为其组成部分(段落、条款)及结构分组(卷、标题、章等)创建嵌入表示,通过密集向量捕捉法律信息的细微差异,实现多层次语义表征。该方法使检索增强生成系统能根据用户查询,准确返回特定片段或完整章节的内容,满足多样化信息需求。研究探讨了法律文本中的相关性、语义切块与固有层级结构,表明该方法显著提升法律信息检索效果。尽管以巴西立法体系和《巴西宪法》(civil law tradition)为实验基础,但原理原则上可适用于普通法体系及其他层级化文本领域。该方法的通用原则亦可拓展至政策、标准等具有层次结构的文本组织与检索场景。

原文摘要 · Abstract (English)

This work addresses the challenge of capturing the complexities of legal knowledge by proposing a multi-layered embedding-based retrieval method for legal and legislative texts. Creating embeddings not only for individual articles but also for their components (paragraphs, clauses) and structural groupings (books, titles, chapters, etc), we seek to capture the subtleties of legal information through the use of dense vectors of embeddings, representing it at varying levels of granularity. Our method meets various information needs by allowing the Retrieval Augmented Generation system to provide accurate responses, whether for specific segments or entire sections, tailored to the user's query. We explore the concepts of aboutness, semantic chunking, and inherent hierarchy within legal texts, arguing that this method enhances the legal information retrieval. Despite the focus being on Brazil's legislative methods and the Brazilian Constitution, which follow a civil law tradition, our findings should in principle be applicable across different legal systems, including those adhering to common law traditions. Furthermore, the principles of the proposed method extend beyond the legal domain, offering valuable insights for organizing and retrieving information in any field characterized by information encoded in hierarchical text.

法律AI信息检索多粒度嵌入层次文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。