首个在人工标注平行语料上评估跨语言话语解析的论文,效果超越现有方法。
Bilingual Rhetorical Structure Parsing with Large Parallel Annotations
- 基于英语GUM语料构建俄语平行标注,实现双语话语结构对齐。
- 端到端模型在英俄语料上均达顶尖性能,跨语言迁移有效。
- 适合研究多语言文本结构分析与跨语言自然语言处理的学者。
话语解析是自然语言处理中的关键任务,旨在揭示文本的高层级语篇关系。尽管跨语言话语解析受到越来越多关注,但受限于平行数据稀少以及不同语言和语料中对修辞结构理论(RST)应用不一致的问题,仍面临挑战。为此,我们为大型多样化的英语GUM RST语料库引入了俄语平行标注。利用最新进展,我们的端到端RST解析器在英语和俄语语料库上均达到当前最优结果。该模型在单语和双语设置下均表现优异,即使第二语言标注数据有限也能实现有效迁移。据我们所知,这是首次在人工标注的平行语料上评估跨语言端到端RST解析潜力的工作。
原文摘要 · Abstract (English)
Discourse parsing is a crucial task in natural language processing that aims to reveal the higher-level relations in a text. Despite growing interest in cross-lingual discourse parsing, challenges persist due to limited parallel data and inconsistencies in the Rhetorical Structure Theory (RST) application across languages and corpora. To address this, we introduce a parallel Russian annotation for the large and diverse English GUM RST corpus. Leveraging recent advances, our end-to-end RST parser achieves state-of-the-art results on both English and Russian corpora. It demonstrates effectiveness in both monolingual and bilingual settings, successfully transferring even with limited second-language annotation. To the best of our knowledge, this work is the first to evaluate the potential of cross-lingual end-to-end RST parsing on a manually annotated parallel corpus.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。