arXiv:2605.26926cs.AI2026-05

用智能代理+检索增强生成,让法律条文自动变可量化的监测指标。

From Norms to Indicators (N2I-RAG): An Agentic Retrieval-Augmented Generation Framework for Legal Indicator Computation

论文配图:From Norms to Indicators (N2I-RAG): An Agentic Retrieval-Augmented Generation Framework for Legal Indicator Computation
图 1 · 摘自论文原文
  • 构建模块化代理系统,分步完成法律条文检索与证据评估。
  • 在法语海洋环保法规上测试,准确率显著高于基线模型。
  • 适合需要透明、可追溯的法律监测与政策评估场景。

从规范性法律文本中计算法律指标是法律监管与政策评估的关键任务,但因法律语言复杂、规模大且需解释,加之文档质量参差,带来巨大挑战。现有自然语言处理与生成模型常出现幻觉,缺乏可解释性与证据支撑,难以保障指标可靠性。本文提出N2I-RAG(From Norms to Indicators),一种面向法律指标计算的智能体检索增强生成框架,通过自适应检索、基于大模型的智能体与验证机制,在模块化流程中实现证据筛选、检索与评估,并输出可关联具体法律条款的二元法律结论。该框架强调可追溯性,要求对中间决策与最终指标赋值提供明确解释。我们在自建的法语海洋环境法语料库(含扫描件与数字源)上进行评估,涵盖多种语言模型家族。对比实验表明,该方法持续优于基线系统,且在两个不同禁令上具有良好的泛化能力。结果表明,智能体检索增强生成能有效连接开放文本的法律语言与标准化指标计算,为透明、可扩展的法律观测体系奠定基础。

原文摘要 · Abstract (English)

Computing legal indicators from normative texts is a key task in legal monitoring and policy evaluation, but presents significant challenges due to the complexity, scale, and interpretive nature of legal language, as well as the variability in available document quality. Existing natural language processing techniques and generative models can assist in legal analysis, but often suffer from high risk of hallucinations and lack the interpretability and evidence grounding required for reliable indicator computation. This paper presents N2I-RAG (From Norms to Indicators), an agentic retrieval-augmented generation framework designed to automate the computation of legal indicators in a transparent and traceable way. We integrate adaptive retrieval, llm-based agents, and validation mechanisms in a modular pipeline, where each component performs a defined role in filtering, retrieving, and assessing evidence, and in producing binary legal outcomes linked to identifiable legal provisions. The framework emphasizes traceability by requiring explicit explanations of intermediate decisions and final indicator assignments. We evaluate N2I-RAG using an in-house constructed French marine environmental law corpus that includes both scanned and digital sources. Comparative experiments with multiple language model families demonstrate that the proposed approach consistently outperforms baseline systems, and generalizes well when tested on 2 different bans. The results indicate that agentic retrieval-augmented generation can bridge open-text legal language and standardized indicator computation, offering a foundation for transparent and scalable legal observatories.

法律AIRAG智能体指标计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。