构建首个高质量葡萄牙语开放信息抽取语料库,解决资源匮乏问题。
Challenges in Expanding Portuguese Resources: A View from Open Information Extraction
- 基于语义理论设计严谨标注规则,确保语料质量
- 验证了现有先进模型在葡萄牙语上的性能表现
- 适合研究多语言NLP与低资源语言信息抽取者
开放信息抽取(Open IE)旨在从文本中无领域依赖地提取结构化信息。尽管近年来数据驱动方法显著提升效果,但主要集中在英语,其他语言因缺乏标注数据而受限。本文提出一个基于严谨语义理论的高质量葡萄牙语手动标注语料库,讨论标注过程中的挑战,制定结构与上下文标注规则,并通过评估前沿Open IE系统验证其有效性。该资源填补了葡萄牙语Open IE数据空白,可支持新方法开发与评估。
原文摘要 · Abstract (English)
Open Information Extraction (Open IE) is the task of extracting structured information from textual documents, independent of domain. While traditional Open IE methods were based on unsupervised approaches, recently, with the emergence of robust annotated datasets, new data-based approaches have been developed to achieve better results. These innovations, however, have focused mainly on the English language due to a lack of datasets and the difficulty of constructing such resources for other languages. In this work, we present a high-quality manually annotated corpus for Open Information Extraction in the Portuguese language, based on a rigorous methodology grounded in established semantic theories. We discuss the challenges encountered in the annotation process, propose a set of structural and contextual annotation rules, and validate our corpus by evaluating the performance of state-of-the-art Open IE systems. Our resource addresses the lack of datasets for Open IE in Portuguese and can support the development and evaluation of new methods and systems in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。