自动标注指代消解数据,降低人工标注成本。
Towards Generating Automatic Anaphora Annotations
- 从现有数据直接转换生成指代标注数据
- 用多语言模型解析新语言的指代关系
- 适合需要大规模指代数据的研究者
训练在各类自然语言处理任务中表现优异的模型需要大量数据,而指代消解这类细粒度任务对数据需求尤为突出。为应对人工标注金标准数据的巨大成本,本文探索两种自动生成带有共指标注数据集的方法:一是直接转换现有数据集,二是利用具备跨语言能力的多语言模型解析新出现或未见语言的指代关系。论文详细介绍了这两方面当前的进展,以及所面临挑战,并提出了相应的应对策略。
原文摘要 · Abstract (English)
Training models that can perform well on various NLP tasks require large amounts of data, and this becomes more apparent with nuanced tasks such as anaphora and conference resolution. To combat the prohibitive costs of creating manual gold annotated data, this paper explores two methods to automatically create datasets with coreferential annotations; direct conversion from existing datasets, and parsing using multilingual models capable of handling new and unseen languages. The paper details the current progress on those two fronts, as well as the challenges the efforts currently face, and our approach to overcoming these challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。