对比三种NER方法在德国行政法规文本分析中的表现,发现判别式模型最优。
GerPS-Compare: Comparing NER methods for legal norm analysis
- 用规则、判别式和生成式三种NER方法分析行政法规文本
- 判别式模型整体优于规则与生成模型,尤其在类别异质性强时表现更好
- 适合法律文本结构复杂、类别差异大的场景,对法务研究者有实用价值
我们将命名实体识别(NER)应用于德语中特定子类型法律文本——规范公共服务行政流程的法律条文。此类文本分析需识别十类由公共行政专业人员定义的文本片段。本文比较了三种NER方法:基于规则的系统、深度判别模型和深度生成模型。结果表明,深度判别模型在整体性能上优于基于规则的系统及深度生成模型;后两者表现相近,但在不同类别上各有优劣。这一出乎意料的结果主要源于所用类别在语义和句法上的高度异质性,不同于常规NER任务中的类别。深度判别模型在此类复杂结构下表现出更强适应性,优于通用大模型和人工设计的规则系统。
原文摘要 · Abstract (English)
We apply NER to a particular sub-genre of legal texts in German: the genre of legal norms regulating administrative processes in public service administration. The analysis of such texts involves identifying stretches of text that instantiate one of ten classes identified by public service administration professionals. We investigate and compare three methods for performing Named Entity Recognition (NER) to detect these classes: a Rule-based system, deep discriminative models, and a deep generative model. Our results show that Deep Discriminative models outperform both the Rule-based system as well as the Deep Generative model, the latter two roughly performing equally well, outperforming each other in different classes. The main cause for this somewhat surprising result is arguably the fact that the classes used in the analysis are semantically and syntactically heterogeneous, in contrast to the classes used in more standard NER tasks. Deep Discriminative models appear to be better equipped for dealing with this heterogenerity than both generic LLMs and human linguists designing rule-based NER systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。