arXiv:2505.04132cs.CLcs.AI2025-05被引 11

用大模型生成法律问答库,让普通人也能读懂法律条文。

Bringing legal knowledge to the public by constructing a legal question bank using large-scale pre-trained language model

  • 用大模型将法律条文转为通俗易懂的短片段(CLIC页)
  • 自动生成的法律问题比人工更高效多样,准确率也高
  • 适合法律科普、公众咨询和司法服务系统开发者

获取法律信息是实现司法公正的基础。然而,可及性不仅指公开法律文件,还在于让普通人理解这些信息。将专业法律文本转化为普通人可读的内容是一大难题。本研究提出三步法:首先将法律条文拆解为解释特定法律概念的通俗片段(称为CLIC页);其次构建法律问答库(LQB),包含可从CLIC页找到答案的问题;最后设计交互式CLIC推荐器(CRec),根据用户对法律情境的描述,推荐最相关的问题与对应CLIC页。本文聚焦于构建LQB的技术实现,展示如何利用GPT-3等大模型生成法律问题。对比机器生成问题(MGQs)与人工编写的問題(HCQs),发现前者在可扩展性、成本效益和多样性上更优,而后者更精准。我们还展示了CRec原型,并通过实例说明该方法能有效将法律知识推送给公众。

原文摘要 · Abstract (English)

Access to legal information is fundamental to access to justice. Yet accessibility refers not only to making legal documents available to the public, but also rendering legal information comprehensible to them. A vexing problem in bringing legal information to the public is how to turn formal legal documents such as legislation and judgments, which are often highly technical, to easily navigable and comprehensible knowledge to those without legal education. In this study, we formulate a three-step approach for bringing legal knowledge to laypersons, tackling the issues of navigability and comprehensibility. First, we translate selected sections of the law into snippets (called CLIC-pages), each being a small piece of article that focuses on explaining certain technical legal concept in layperson's terms. Second, we construct a Legal Question Bank (LQB), which is a collection of legal questions whose answers can be found in the CLIC-pages. Third, we design an interactive CLIC Recommender (CRec). Given a user's verbal description of a legal situation that requires a legal solution, CRec interprets the user's input and shortlists questions from the question bank that are most likely relevant to the given legal situation and recommends their corresponding CLIC pages where relevant legal knowledge can be found. In this paper we focus on the technical aspects of creating an LQB. We show how large-scale pre-trained language models, such as GPT-3, can be used to generate legal questions. We compare machine-generated questions (MGQs) against human-composed questions (HCQs) and find that MGQs are more scalable, cost-effective, and more diversified, while HCQs are more precise. We also show a prototype of CRec and illustrate through an example how our 3-step approach effectively brings relevant legal knowledge to the public.

法律AI大模型应用法律科普

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。