arXiv:2410.08521cs.CLcs.IR2024-10被引 4

用语义过滤提升法律实体识别准确率,处理法律文本中的模糊与嵌套问题。

Improving Legal Entity Recognition Using a Hybrid Transformer Model and Semantic Filtering Approach

  • 融合Transformer与语义相似度过滤,增强模型对法律文本的理解能力。
  • 在1.5万份标注文档上达到93.4%的F1分数,精度与召回率均显著提升。
  • 适合法律科技、智能合同分析等需要高精度实体识别的场景。

法律实体识别(LER)在自动化合同分析、合规监控和诉讼支持等法律流程中至关重要。现有方法如规则系统和传统机器学习模型难以应对法律文档的复杂性与领域特异性,尤其在处理歧义和嵌套实体结构时表现不佳。本文提出一种新型混合模型,通过引入基于语义相似度的过滤机制,提升微调后的Legal-BERT模型在法律文本处理中的表现。我们在包含15,000份标注法律文档的数据集上进行评估,取得了93.4%的F1分数,相比先前方法在精度和召回率上均有显著提升。

原文摘要 · Abstract (English)

Legal Entity Recognition (LER) is critical in automating legal workflows such as contract analysis, compliance monitoring, and litigation support. Existing approaches, including rule-based systems and classical machine learning models, struggle with the complexity of legal documents and domain specificity, particularly in handling ambiguities and nested entity structures. This paper proposes a novel hybrid model that enhances the accuracy and precision of Legal-BERT, a transformer model fine-tuned for legal text processing, by introducing a semantic similarity-based filtering mechanism. We evaluate the model on a dataset of 15,000 annotated legal documents, achieving an F1 score of 93.4%, demonstrating significant improvements in precision and recall over previous methods.

法律AI实体识别Transformer语义过滤

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。