arXiv:2510.06427cs.CL2025-10中稿 · CODI CRAC 2025

首个统一多语言修辞结构解析器,支持18个语料库无需修改标签体系。

Bridging Discourse Treebanks with a Unified Rhetorical Structure Parser

  • 采用分头分类与标签掩码共享参数两种训练策略应对标签体系差异。
  • 在18个语料库上表现优于16个单语基线模型,验证统一模型优势。
  • 适合需要跨语言、多语料对话分析的研究者使用。

我们提出UniRST,首个无需修改11种语言中18个语料库关系标注体系的统一式修辞结构(RST)解析器。为克服标注体系不兼容问题,我们设计并评估了两种训练策略:多头(Multi-Head)为每套标注体系分配独立分类层,掩码联合(Masked-Union)通过选择性标签掩码实现参数共享。首先在低资源场景下引入一种简单有效的数据增强技术以提升单语语料库解析性能。随后训练统一模型,结果表明:(1)参数高效的掩码联合方法同时达到最佳性能;(2)UniRST在18个语料库中有16个超过对应单语基线,证明单一模型、端到端多语言对话解析在多样化资源上的优越性。

原文摘要 · Abstract (English)

We introduce UniRST, the first unified RST-style discourse parser capable of handling 18 treebanks in 11 languages without modifying their relation inventories. To overcome inventory incompatibilities, we propose and evaluate two training strategies: Multi-Head, which assigns separate relation classification layer per inventory, and Masked-Union, which enables shared parameter training through selective label masking. We first benchmark monotreebank parsing with a simple yet effective augmentation technique for low-resource settings. We then train a unified model and show that (1) the parameter efficient Masked-Union approach is also the strongest, and (2) UniRST outperforms 16 of 18 mono-treebank baselines, demonstrating the advantages of a single-model, multilingual end-to-end discourse parsing across diverse resources.

对话分析多语言修辞结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。