arXiv:2607.11109cs.IRcs.CL2026-07

用生成模型让法律查询更懂人话,精准匹配法条

Generative Chinese Statute Retrieval

  • 将法条检索转为文本生成任务,内化法律知识
  • 多粒度文档编码+多任务训练,显著提升检索效果
  • 适合法律AI研发者和智能司法系统构建者

法条检索是法律信息检索的基础任务,但现有方法难以弥合口语化法律查询与正式法条语言之间的差距。本文提出GCSR,一种生成式法条检索框架,将法条检索重构为序列生成问题,并将法条知识内化至生成模型中。具体而言,我们设计了多粒度结构化docid,以编码法律层级与语义信息,并采用多任务训练策略。实验表明,GCSR在多个基准上持续优于强大多样检索基线(包括稀疏、稠密及法律领域模型)。结果验证了生成式检索在法条检索中的有效性,凸显其在更广泛的法律信息获取与下游法律推理任务中的潜力。

原文摘要 · Abstract (English)

Statute retrieval is a fundamental task in legal information retrieval, yet existing approaches struggle to bridge the gap between colloquial legal queries and formal statutory language. In this paper, we propose GCSR, a generative statute retrieval framework that reformulates statute retrieval as a sequence generation problem and internalizes statutory knowledge into a generative model. Specifically, we propose a multi-granularity structured docid that encodes legal hierarchy and semantic information, together with a multi-task training strategy. Experiments show that GCSR consistently outperforms strong sparse, dense, and legal-domain baselines. Our results demonstrate the effectiveness of generative retrieval for statute retrieval and highlight its potential for broader legal information access and downstream legal reasoning tasks.

法律AI生成检索法条匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。