构建法律问答数据集,让AI用少量权威文献精准回答普通人法律问题。
Experimenting with Legal AI Solutions: The Case of Question-Answering for Access to Justice
- 设计面向普通人的法律NLP流程,涵盖数据采集、推理与评估
- 仅用850条引用文献的检索增强生成,效果媲美全网检索
- 开源真实法律问答数据集,推动开放模型发展
生成式AI模型(如GPT和Llama系列)在帮助普通人解答法律问题方面具有巨大潜力。然而,以往研究较少关注面向普通用户的数据获取、推理与评估。为此,我们提出一个以人类为中心的法律自然语言处理流程,涵盖数据采集、推理与评估。我们引入并发布了一个名为LegalQA的数据集,包含从劳动法到刑事法等领域的实际具体法律问题,由法律专家撰写的对应答案,以及每条答案的引用来源。我们开发了针对该数据集的自动评估协议,并表明:仅使用训练集中850条引用文献进行检索增强生成,即可达到或超越互联网范围检索的效果,尽管数据量少9个数量级。最后,我们提出了未来开源努力的方向,指出当前开源模型仍落后于闭源模型。
原文摘要 · Abstract (English)
Generative AI models, such as the GPT and Llama series, have significant potential to assist laypeople in answering legal questions. However, little prior work focuses on the data sourcing, inference, and evaluation of these models in the context of laypersons. To this end, we propose a human-centric legal NLP pipeline, covering data sourcing, inference, and evaluation. We introduce and release a dataset, LegalQA, with real and specific legal questions spanning from employment law to criminal law, corresponding answers written by legal experts, and citations for each answer. We develop an automatic evaluation protocol for this dataset, then show that retrieval-augmented generation from only 850 citations in the train set can match or outperform internet-wide retrieval, despite containing 9 orders of magnitude less data. Finally, we propose future directions for open-sourced efforts, which fall behind closed-sourced models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。