首个中文法律诉求生成数据集,助力普通人写诉状
ClaimGen-CN: A Large-scale Chinese Dataset for Legal Claim Generation
- 构建首个中文法律诉求生成数据集ClaimGen-CN
- 现有模型在事实准确性和表达清晰度上表现不佳
- 适合法律AI研究者和司法科技开发者使用
法律诉求是原告在案件中的请求,对司法推理和案件解决至关重要。尽管已有大量研究致力于提升法律专业人士的效率,但帮助非专业人士(如原告)生成诉求的研究仍属空白。本文探讨基于案件事实生成法律诉求的问题。首先,我们从真实法律纠纷中构建了ClaimGen-CN,这是首个面向中文法律诉求生成任务的数据集。此外,我们设计了专门的评估指标,涵盖事实性和清晰度两个核心维度。基于此,我们对当前最先进的通用及法律领域大语言模型进行了全面的零样本评估。结果表明,现有模型在事实精确性和表达清晰度方面存在明显不足,凸显出该领域亟需针对性改进。为推动该重要任务的发展,我们将公开发布该数据集。
原文摘要 · Abstract (English)
Legal claims refer to the plaintiff's demands in a case and are essential to guiding judicial reasoning and case resolution. While many works have focused on improving the efficiency of legal professionals, the research on helping non-professionals (e.g., plaintiffs) remains unexplored. This paper explores the problem of legal claim generation based on the given case's facts. First, we construct ClaimGen-CN, the first dataset for Chinese legal claim generation task, from various real-world legal disputes. Additionally, we design an evaluation metric tailored for assessing the generated claims, which encompasses two essential dimensions: factuality and clarity. Building on this, we conduct a comprehensive zero-shot evaluation of state-of-the-art general and legal-domain large language models. Our findings highlight the limitations of the current models in factual precision and expressive clarity, pointing to the need for more targeted development in this domain. To encourage further exploration of this important task, we will make the dataset publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。