剖析预训练模型在隐式连贯关系识别中的表现,提升低数据场景下的识别效果。
Fine-Grained Evaluation for Implicit Discourse Relation Recognition
- 通过深度分析模型预测,揭示预训练模型在该任务中的难点。
- 半自动标注新数据,显著提升二级语义关系的识别效果。
- 适合关注细粒度话语分析与数据稀缺问题的研究者。
隐式话语关系识别因缺乏显式连接词而具有挑战性。近年来预训练语言模型在此任务上取得显著进展,但对其性能缺乏细粒度分析,导致任务难点和潜在方向不明确。本文深入分析预训练模型的预测结果,尝试揭示其困难所在及可能改进方向。此外,为提升标注数据稀缺的关系类型表现,我们基于PDTB 3.0中样本较少的类别进行半自动标注,构建高质量补充数据。实验表明,新增数据显著改善了二级语义关系的识别性能。
原文摘要 · Abstract (English)
Implicit discourse relation recognition is a challenging task in discourse analysis due to the absence of explicit discourse connectives between spans of text. Recent pre-trained language models have achieved great success on this task. However, there is no fine-grained analysis of the performance of these pre-trained language models for this task. Therefore, the difficulty and possible directions of this task is unclear. In this paper, we deeply analyze the model prediction, attempting to find out the difficulty for the pre-trained language models and the possible directions of this task. In addition to having an in-depth analysis for this task by using pre-trained language models, we semi-manually annotate data to add relatively high-quality data for the relations with few annotated examples in PDTB 3.0. The annotated data significantly help improve implicit discourse relation recognition for level-2 senses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。