arXiv:2508.19365cs.IRcs.CY2025-08被引 5

用真实法律数据评估AI简化法规的可靠性,发现效果远未达宣传水平。

AI for Statutory Simplification: A Comprehensive State Legal Corpus and Labor Benchmark

  • 构建劳动法对比数据集LaborBench,覆盖50州101维差异
  • 8.7GB州级法规语料库+问答测试,验证AI提取准确性
  • 大模型在复杂法规处理中仍存显著错误,不适合作为最终简化工具

AI在法律领域的新兴应用之一是法规简化:对复杂的立法或监管语言进行梳理与精简。美国一州宣称已通过AI消除其州法三分之一内容。然而,此类方法的准确性、可靠性及风险尚无系统评估。本文提出LaborBench——一个面向问答的基准数据集,用于评估AI在此领域的能力。该数据源自美国劳工部(DOL)每年由律师团队历时六个月编制的更新数据,涵盖50个州在超过101个维度上的失业保险法差异,最终形成200页的表格出版物。受与一州合作探索使用大语言模型(LLMs)简化法规的启发,我们将此出版物转化为LaborBench,提供评估AI进行法规信息提炼与提取能力的独特基准。为支持系统性研究,我们还构建了StateCodes——一个全新且全面的州级法规语料库,共8.7 GB。我们对信息检索与前沿大模型在该数据集上的表现进行基准测试,结果表明:尽管这些模型可作为初步研究辅助,但整体准确率远低于对大模型端到端简化监管文本所宣称的预期。

原文摘要 · Abstract (English)

One of the emerging use cases of AI in law is for code simplification: streamlining, distilling, and simplifying complex statutory or regulatory language. One U.S. state has claimed to eliminate one third of its state code using AI. Yet we lack systematic evaluations of the accuracy, reliability, and risks of such approaches. We introduce LaborBench, a question-and-answer benchmark dataset designed to evaluate AI capabilities in this domain. We leverage a unique data source to create LaborBench: a dataset updated annually by teams of lawyers at the U.S. Department of Labor, who compile differences in unemployment insurance laws across 50 states for over 101 dimensions in a six-month process, culminating in a 200-page publication of tables. Inspired by our collaboration with one U.S. state to explore using large language models (LLMs) to simplify codes in this domain, where complexity is particularly acute, we transform the DOL publication into LaborBench. This provides a unique benchmark for AI capacity to conduct, distill, and extract realistic statutory and regulatory information. To assess the performance of retrieval augmented generation (RAG) approaches, we also compile StateCodes, a novel and comprehensive state statute and regulatory corpus of 8.7 GB, enabling much more systematic research into state codes. We then benchmark the performance of information retrieval and state-of-the-art large LLMs on this data and show that while these models are helpful as preliminary research for code simplification, the overall accuracy is far below the touted promises for LLMs as end-to-end pipelines for regulatory simplification.

法律AI大模型评测法规简化数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。