通过对比局部语法差异,选出最适合提取葡萄牙语人名的最优语法
Concordance Comparison as a Means of Assembling Local Grammars

- 用共现分析法比较两组局部语法的异同
- 在葡萄牙语人名识别上达76.86的F值,提升6点
- 适合需要精准命名实体识别的多语言研究者
人名识别是信息抽取中的重要但复杂任务。本文提出一种方法,通过比较两个局部语法(LG)所得的共现频次,识别其包含、交集与互斥关系,据此筛选出最优语法组合。该方法在葡萄牙语文本中进行案例研究,应用于第二届HAREM竞赛的Gold Collection数据集,最终取得76.86的F-Measure,较现有最佳水平提升6个百分点。
原文摘要 · Abstract (English)
Named Entity Recognition for person names is an important but non-trivial task in information extraction. This article uses a tool that compares the concordances obtained from two local grammars (LG) and highlights the differences. We used the results as an aid to select the best of a set of LGs. By analyzing the comparisons, we observed relationships of inclusion, intersection and disjunction within each pair of LGs, which helped us to assemble those that yielded the best results. This approach was used in a case study on extraction of person names from texts written in Portuguese. We applied the enhanced grammar to the Gold Collection of the Second HAREM. The F-Measure obtained was 76.86, representing a gain of 6 points in relation to the state-of-the-art for Portuguese.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。