构建首个跨文体开放数据集,提升英文浅层语篇解析泛化能力
GDTB: Genre Diverse Data for English Shallow Discourse Parsing across Modalities, Text Types, and Domains
- 基于UD GUM语料构建多文体、跨领域新数据集
- 跨领域测试中性能下降明显,联合训练可缓解此问题
- 适合研究跨域泛化与语篇分析的学者使用
英语浅层语篇解析研究长期依赖《华尔街日报》语料,但该数据集已封闭、仅限新闻领域且年代久远。本文基于已有标注的UD英语GUM语料,构建并评估了一个新的开源、多文体基准数据集,用于PDTB风格的浅层语篇解析。在一系列跨领域关系分类实验中发现,尽管新数据集与PDTB兼容,但存在显著的域外性能下降;通过联合训练两个数据集可有效缓解该问题。
原文摘要 · Abstract (English)
Work on shallow discourse parsing in English has focused on the Wall Street Journal corpus, the only large-scale dataset for the language in the PDTB framework. However, the data is not openly available, is restricted to the news domain, and is by now 35 years old. In this paper, we present and evaluate a new open-access, multi-genre benchmark for PDTB-style shallow discourse parsing, based on the existing UD English GUM corpus, for which discourse relation annotations in other frameworks already exist. In a series of experiments on cross-domain relation classification, we show that while our dataset is compatible with PDTB, substantial out-of-domain degradation is observed, which can be alleviated by joint training on both datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。