arXiv:2602.00881cs.CL2026-02中稿 · Publication at EAC…

构建印度法律条文识别语料库,解决普通人提问与法院文书差异问题

ILSIC: Corpora for Identifying Indian Legal Statutes from Queries by Laypeople

  • 构建包含500+条印度法律条文的普通人提问语料库ILSIC
  • 模型仅用法院判例训练时在普通人查询上表现差,迁移学习可提升效果
  • 提供法院判决数据用于对比分析,适合法律NLP与司法智能化研究者

法律条文识别(LSI)是法律自然语言处理中的基础任务。传统上该任务依赖法院判例中的事实作为输入查询,因数据丰富。但在实际应用中,查询往往来自非专业人士且表达非正式。尽管已有少数面向普通人的LSI数据集,但针对法院与普通人数据差异的研究仍较少。本文构建了ILSIC,一个涵盖500+条印度法律条文的普通人查询语料库,并附带相应法院判决,便于研究者比较两类数据在LSI上的表现。我们在该语料库上进行了广泛实验,包括零样本、少样本推理、检索增强生成及监督微调。结果表明:仅在法院判例上训练的模型在普通人查询上表现不佳;而从法院数据向普通人数据进行迁移学习在某些场景下有显著收益。此外,我们还对查询类别和条文频次进行了细粒度分析。

原文摘要 · Abstract (English)

Legal Statute Identification (LSI) for a given situation is one of the most fundamental tasks in Legal NLP. This task has traditionally been modeled using facts from court judgments as input queries, due to their abundance. However, in practical settings, the input queries are likely to be informal and asked by laypersons, or non-professionals. While a few laypeople LSI datasets exist, there has been little research to explore the differences between court and laypeople data for LSI. In this work, we create ILSIC, a corpus of laypeople queries covering 500+ statutes from Indian law. Additionally, the corpus also contains court case judgements to enable researchers to effectively compare between court and laypeople data for LSI. We conducted extensive experiments on our corpus, including benchmarking over the laypeople dataset using zero and few-shot inference, retrieval-augmented generation and supervised fine-tuning. We observe that models trained purely on court judgements are ineffective during test on laypeople queries, while transfer learning from court to laypeople data can be beneficial in certain scenarios. We also conducted fine-grained analyses of our results in terms of categories of queries and frequency of statutes.

法律AI语料库自然语言处理迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。