arXiv:2410.12064cs.CLcs.LG2024-10被引 10

聚焦法律文本中的违规识别,提升法律条文与责任人关联的自动化能力。

LegalLens Shared Task 2024: Legal Violation Identification in Unstructured Text

  • 基于预训练模型微调,融合多领域法律文本提升识别效果。
  • NER任务最佳模型较基线提升7.11%,NLI任务提升5.7%。
  • 适合法律AI研究者与合规科技开发者参考。

本文报告了LegalLens共享任务2024年的结果,重点在于从非结构化文本中检测法律违规行为,涵盖两个子任务:LegalLens-NER用于识别法律违规实体,LegalLens-NLI用于将这些违规行为与相关法律条文及受影响个体关联。使用覆盖劳动、隐私和消费者保护领域的增强版LegalLens数据集,共有38支队伍参与。分析显示,尽管方法多样,但两任务中表现最优的团队均依赖预训练语言模型的微调,优于专用法律模型和少样本方法。最佳团队在NER任务上相较基线提升7.11%,而NLI任务提升为5.7%。尽管取得进展,法律文本的复杂性仍表明存在进一步优化空间。

原文摘要 · Abstract (English)

This paper presents the results of the LegalLens Shared Task, focusing on detecting legal violations within text in the wild across two sub-tasks: LegalLens-NER for identifying legal violation entities and LegalLens-NLI for associating these violations with relevant legal contexts and affected individuals. Using an enhanced LegalLens dataset covering labor, privacy, and consumer protection domains, 38 teams participated in the task. Our analysis reveals that while a mix of approaches was used, the top-performing teams in both tasks consistently relied on fine-tuning pre-trained language models, outperforming legal-specific models and few-shot methods. The top-performing team achieved a 7.11% improvement in NER over the baseline, while NLI saw a more marginal improvement of 5.7%. Despite these gains, the complexity of legal texts leaves room for further advancements.

法律AI实体识别文本理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。