用AI整合法律知识图谱与向量库,提升法律文本检索与推理准确率。
Bridging Legal Knowledge and AI: Retrieval-Augmented Generation with Vector Stores, Knowledge Graphs, and Hierarchical Non-negative Matrix Factorization
- 融合RAG、向量库与知识图谱,通过非负矩阵分解构建层级关系
- 在公开法律数据上实现跨文档关联分析,支持法律条文与判例的智能推理
- 适合法律AI研究者与司法智能化开发者,减少幻觉并提升可解释性
基于大语言模型的生成式代理AI,结合检索增强生成(RAG)、知识图谱(KG)和向量存储(VS),在法律等专业领域具有变革性应用潜力。该技术擅长从海量非结构化或半结构化数据中推断复杂关系。本文针对法律系统中宪法、法规、规章与判例构成的复杂、互联、半结构化知识体系,提出一种集成RAG、VS与基于非负矩阵分解(NMF)构建的知识图谱的生成式AI系统。该系统利用网络爬虫从Justia等公开平台系统收集法律文本,通过高级语义表征、层级关系与潜在主题发现,突破传统关键词搜索局限,实现法律文件聚类、摘要与交叉引用。系统显著提升法律信息检索的可扩展性、可解释性与准确性,助力计算法学发展,并支持对法律案例、法规间复杂关联的识别与法律趋势预测,有效降低幻觉风险。
原文摘要 · Abstract (English)
Agentic Generative AI, powered by Large Language Models (LLMs) with Retrieval-Augmented Generation (RAG), Knowledge Graphs (KGs), and Vector Stores (VSs), represents a transformative technology applicable to specialized domains such as legal systems, research, recommender systems, cybersecurity, and global security, including proliferation research. This technology excels at inferring relationships within vast unstructured or semi-structured datasets. The legal domain here comprises complex data characterized by extensive, interrelated, and semi-structured knowledge systems with complex relations. It comprises constitutions, statutes, regulations, and case law. Extracting insights and navigating the intricate networks of legal documents and their relations is crucial for effective legal research. Here, we introduce a generative AI system that integrates RAG, VS, and KG, constructed via Non-Negative Matrix Factorization (NMF), to enhance legal information retrieval and AI reasoning and minimize hallucinations. In the legal system, these technologies empower AI agents to identify and analyze complex connections among cases, statutes, and legal precedents, uncovering hidden relationships and predicting legal trends-challenging tasks that are essential for ensuring justice and improving operational efficiency. Our system employs web scraping techniques to systematically collect legal texts, such as statutes, constitutional provisions, and case law, from publicly accessible platforms like Justia. It bridges the gap between traditional keyword-based searches and contextual understanding by leveraging advanced semantic representations, hierarchical relationships, and latent topic discovery. This framework supports legal document clustering, summarization, and cross-referencing, for scalable, interpretable, and accurate retrieval for semi-structured data while advancing computational law and AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。